Evals and Shipping · Measuring · lesson 1 of 7
Why 'it looked fine' is not a test
about 16 minutes · free · runs in your browser
No compiler, no type system, no red squiggle
Change a function signature and the compiler finds every caller. Change a prompt and nothing happens at all — until a week later, when someone notices the extraction has been returning nulls for a whole category of input.
The reason is worth stating plainly: an LLM feature has no failing build. Its regressions are silent by construction, so they have to be measured deliberately or not at all.
A golden dataset is the smallest thing that works: inputs paired with what a correct answer must contain. Twenty examples beat none by an enormous margin, and the first ten take twenty minutes — the hard part is that nobody sits down to write them.
GOLDEN = [
{"input": "the delivery was fast and excellent", "expect": "positive"},
{"input": "it arrived broken and I want a refund", "expect": "negative"},
]
Your turn: write evaluate(cases, fn) returning (passed, failures) — how many
cases passed, and the list of inputs that did not. The failing inputs are the point: a
score with no examples tells you something is wrong and not what.
You start from this, and edit it in the browser:
def evaluate(cases, fn):
"""Run fn on every case. Return (passed, failing_inputs)."""
return (0, [])