Evals and Shipping · Measuring · lesson 3 of 7
LLM as judge, carefully
about 20 minutes · free · runs in your browser
A judge is a measurement instrument, and it drifts
Some questions have no deterministic check. "Is this answer faithful to the source?" "Is the tone right for a support reply?" For those, a second model can grade — but it is a noisy instrument and has to be treated like one.
Four rules that make a judge usable:
- Temperature 0. A grader that disagrees with itself between runs measures nothing.
- A rubric, not an opinion. "Rate 1–5" produces mush. "Answer PASS or FAIL: does every claim appear in the source?" produces something you can count.
- One question at a time. A judge asked to assess accuracy, tone and length at once averages them into a number that means nothing.
- Spot-check the judge. Grade fifty of its verdicts by hand. A judge nobody has audited is a metric nobody should trust.
And the failure that catches everyone: never judge with the prompt being tested. The same model with the same instructions is disposed to like its own output, so the score goes up while nothing improves.
Your turn: write judge(answer, source) returning True for PASS. Ask for exactly
one word, at temperature 0, and treat anything that is not a clean PASS as a fail — an
unparseable verdict is not a pass.
You start from this, and edit it in the browser:
import fake_llm
RUBRIC = "Say something about the answer."
def judge(answer, source):
"""Return True when the judge says the answer is supported by the source."""
return True