Evals and Shipping · Shipping · lesson 7 of 7
Tracing and error taxonomies
about 20 minutes · free · runs in your browser
Step 1 of 2
Explaining a bad answer three weeks later
"The chatbot said something wrong on Tuesday" is a support ticket you can only answer with a trace. For each request, keep enough to reconstruct it:
| Field | Why |
|---|---|
| request id | to tie the steps together |
| model, temperature | the answer depends on both |
| the messages sent | the prompting bug is almost always visible here |
| retrieved chunk ids | to distinguish retrieval failure from generation failure |
| finish reason, usage | truncation and cost |
| latency | to find the slow step, not the slow feature |
Never log the API key, and think hard before logging the full prompt — it contains whatever the user typed, which on a support tool is frequently personal. Log chunk ids rather than chunk text and most of the problem goes away.
Then classify the failures. A count of "errors" is not actionable; a count by kind tells you what to fix first.
Your turn: write classify(record) returning one of "rate_limited",
"timeout", "truncated", "ungrounded", "invalid_output" or "ok", checking
in that order — the earliest cause is the one that explains the rest.
You start from this, and edit it in the browser:
def classify(record):
"""Return the kind of failure this record represents, or 'ok'."""
return "ok"
Step 2 of 2
What to fix first
With every request classified, the summary writes itself — and it is the difference between "the feature is flaky" and "62% of failures are truncation, so raise max_tokens".
The number that matters is the share of failures by kind, not the raw count. Raw counts move with traffic; shares tell you where the problem is.
Your turn: write summarise(records) returning a dictionary mapping each failure
kind to its share of the failures, rounded to two places. "ok" records are excluded —
they are not failures, and including them makes every share look reassuringly small.
You start from this, and edit it in the browser:
def classify(record):
if record.get("status") == 429:
return "rate_limited"
if record.get("timed_out"):
return "timeout"
if record.get("finish_reason") == "length":
return "truncated"
if not record.get("citations"):
return "ungrounded"
if not record.get("parsed"):
():
{}