Evals and Shipping
Why this exists
An LLM feature has no compiler. A prompt edit that breaks an extraction produces no error, no failing build, and no signal at all until someone notices bad rows. Every regression is silent by construction.
Golden datasets
GOLDEN = [
{"input": "the delivery was fast and excellent", "expect": "positive"},
{"input": "it arrived broken and I want a refund", "expect": "negative"},
]
Twenty examples beat none by an enormous margin. Return the failing inputs, not just a score: "0.78" starts an investigation, "these four failed" ends one. A crashing case is a failure, not the end of the run.
Assertions before judges
| Check | How | Cost |
|---|---|---|
| valid JSON | json.loads | free |
| required fields | schema check | free |
| allowed label | set membership | free |
| citation exists | range check | free |
| forbidden phrase | substring | free |
| tone, faithfulness | a judge model | slow, paid, noisy |
Reach for the judge last, and only once the free checks are green.
Using a judge
- Temperature 0 — a grader that disagrees with itself measures noise.
- A rubric, not an opinion — "PASS or FAIL: does every claim appear in the source?" beats "rate 1–5".
- One question at a time — accuracy, tone and length averaged together mean nothing.
- Audit it — grade fifty verdicts by hand.
- Never judge with the prompt being tested — it likes its own output.
An unreadable verdict is a fail, never a pass.
The CI gate
floor = baseline - tolerance
if score < floor:
fail("score " + str(score) + " below floor " + str(floor))
Commit the baseline like any other expectation. Zero tolerance gets the suite disabled within a month; too much tolerance lets a real regression through. Pick it from observed run-to-run variation.
Caching
key = json.dumps({"model": m, "temperature": t, "messages": msgs}, sort_keys=True)
Everything that changes the answer goes in the key. At temperature 0 an exact cache is free correctness. Semantic caching is a real trade: "how do I cancel" and "how do I not cancel" are near neighbours that want opposite answers.
Latency and routing
- Streaming does not shorten the response; it shortens the wait before the first word.
- Route by cheap, legible rules: short and simple → small model; long context or hard task → large.
- Escalate on failure — try small, validate, retry large. Cheaper than always-large and more accurate than always-small.
Tracing
| Keep | Never keep |
|---|---|
| request id, model, temperature | the API key |
| messages sent (with care) | more personal data than you need |
| retrieved chunk ids | chunk text, usually |
| finish reason, usage, latency |
Error taxonomy
if status == 429: return "rate_limited"
if timed_out: return "timeout"
if finish_reason == "length": return "truncated"
if not citations: return "ungrounded"
if not parsed: return "invalid_output"
return "ok"
Order matters. A rate-limited request also has no citations; reporting it as ungrounded sends someone to fix retrieval that never ran.
Report shares of failures, not raw counts. Counts move with traffic; "62% of failures are truncation" comes with an action attached.