Evals and Shipping · Measuring · lesson 2 of 7
Assertions before judges
about 18 minutes · free · runs in your browser
Most of what you care about needs no model to check
The instinct on hearing "evaluate an LLM" is to reach for another LLM. Resist it for as long as possible, because a deterministic assertion is free, instant, and cannot itself be wrong:
| Check | How |
|---|---|
| is it valid JSON? | json.loads |
| does it have the required fields? | schema check |
| is the label one of three? | set membership |
| does it cite a real passage? | range check |
| does it avoid a forbidden phrase? | substring |
| is it short enough? | len |
Every one runs in microseconds and has no API bill. A judge model is for the residue — tone, faithfulness, whether an answer actually addresses the question — and even then only after the cheap checks are green.
Your turn: write assertions(answer, required, forbidden, max_words) returning the
list of failed check names, in this order: "missing:<field>" for each required word not
present, "forbidden:<word>" for each forbidden one that is, and "too_long" when the
answer exceeds max_words. An empty list means everything passed.
You start from this, and edit it in the browser:
def assertions(answer, required, forbidden, max_words):
"""Return the names of the checks that failed, in order."""
return []