Members only

Build an LLM-as-Judge: Automated Evals

Sep18
Friday, September 18, 2026
9:30 AM – 10:15 AM (America / Los Angeles)
Log in to watch

Need an account? Create one free

Facilitated by

Misha Singh
Misha Singh
AI Product Leader at Meta

About Event

Build an LLM-as-Judge: Automated Evals

Most people check if an AI output looks right by reading it themselves. Misha Singh builds the thing that checks it for her.

Misha Singh, AI Product Leader, ex-Meta, starts with a real prompt and a real output, then builds an eval that grades it automatically: one LLM scoring another's work against a rubric she writes live. You'll watch the judge get built step by step, test it against outputs you can already tell are good or bad, and see where it agrees with you and where it doesn't. One working eval, built and tested, screen shared the entire time.

No theory. You follow along and leave with something you can build yourself, using a free-tier LLM you already have access to.

What you'll walk away with

  1. How to write a rubric an LLM can actually grade against. Vague criteria produce vague scores. You'll see how she turns "good output" into checkable conditions a judge model can apply consistently.
  2. Where LLM judges agree with humans, and where they quietly don't. Automated grading carries its own bias. You'll watch her test the judge against cases where it gets confidently wrong.
  3. A pattern you can rebuild today with a free-tier model. No custom infrastructure, no paid eval platform, just a prompt, a rubric, and a loop you can run again on your own project this week.