Skip to content

Evals and Shipping

Why this exists

An LLM feature has no compiler. A prompt edit that breaks an extraction produces no error, no failing build, and no signal at all until someone notices bad rows. Every regression is silent by construction.

Golden datasets

GOLDEN = [
    {"input": "the delivery was fast and excellent", "expect": "positive"},
    {"input": "it arrived broken and I want a refund", "expect": "negative"},
]

Twenty examples beat none by an enormous margin. Return the failing inputs, not just a score: "0.78" starts an investigation, "these four failed" ends one. A crashing case is a failure, not the end of the run.

Assertions before judges

CheckHowCost
valid JSONjson.loadsfree
required fieldsschema checkfree
allowed labelset membershipfree
citation existsrange checkfree
forbidden phrasesubstringfree
tone, faithfulnessa judge modelslow, paid, noisy

Reach for the judge last, and only once the free checks are green.

Using a judge

  1. Temperature 0 — a grader that disagrees with itself measures noise.
  2. A rubric, not an opinion — "PASS or FAIL: does every claim appear in the source?" beats "rate 1–5".
  3. One question at a time — accuracy, tone and length averaged together mean nothing.
  4. Audit it — grade fifty verdicts by hand.
  5. Never judge with the prompt being tested — it likes its own output.

An unreadable verdict is a fail, never a pass.

The CI gate

floor = baseline - tolerance
if score < floor:
    fail("score " + str(score) + " below floor " + str(floor))

Commit the baseline like any other expectation. Zero tolerance gets the suite disabled within a month; too much tolerance lets a real regression through. Pick it from observed run-to-run variation.

Caching

key = json.dumps({"model": m, "temperature": t, "messages": msgs}, sort_keys=True)

Everything that changes the answer goes in the key. At temperature 0 an exact cache is free correctness. Semantic caching is a real trade: "how do I cancel" and "how do I not cancel" are near neighbours that want opposite answers.

Latency and routing

  • Streaming does not shorten the response; it shortens the wait before the first word.
  • Route by cheap, legible rules: short and simple → small model; long context or hard task → large.
  • Escalate on failure — try small, validate, retry large. Cheaper than always-large and more accurate than always-small.

Tracing

KeepNever keep
request id, model, temperaturethe API key
messages sent (with care)more personal data than you need
retrieved chunk idschunk text, usually
finish reason, usage, latency

Error taxonomy

if status == 429:               return "rate_limited"
if timed_out:                   return "timeout"
if finish_reason == "length":   return "truncated"
if not citations:               return "ungrounded"
if not parsed:                  return "invalid_output"
return "ok"

Order matters. A rate-limited request also has no citations; reporting it as ungrounded sends someone to fix retrieval that never ran.

Report shares of failures, not raw counts. Counts move with traffic; "62% of failures are truncation" comes with an action attached.

This cheatsheet is the summary. If you want to build it yourself, the Evals and Shipping course walks you through it in the browser — the first lesson is free.