Structured Output and Tool Calling · Contracts · lesson 3 of 6
Repairing what comes back broken
about 20 minutes · free · runs in your browser
Step 1 of 2
The expensive failure
There is a failure worse than an error: the call succeeds, the tokens are billed, and the body will not parse.
Sure! Here you go:
{"topic": "delivery",}
Two problems in three lines — a chatty preamble the instruction asked it not to write, and a trailing comma JSON does not allow. The HTTP status is 200 and the usage is real.
The instinct is to patch the string: strip everything before the first {, delete
trailing commas with a regex. Resist it. Every patch is a guess about a shape you do not
control, and the regex that fixes today's break mangles tomorrow's valid answer.
The repair that holds is asking again. Show the model what it produced and what was wrong with it, and let it fix its own output. That is a normal call — same API, one more round trip — and it succeeds far more often than string surgery.
fake_llm.break_next("malformed_json") breaks exactly one call, so a second attempt
succeeds — which is what makes this testable.
Your turn: write ask_json_repaired(question, attempts=2) that parses if it can and
re-asks with a repair prompt if it cannot.
You start from this, and edit it in the browser:
import json
import fake_llm
RULE = "Reply with JSON only. No preamble."
def ask_json_repaired(question, attempts=2):
"""Ask for JSON, re-asking with the error when it will not parse."""
response = fake_llm.chat(
[{"role": "system", "content": RULE}, {"role": "user", "content": question}]
)
return json.loads(response["choices"][0]["message"]["content"])
Step 2 of 2
When repair fails too
Retries end. The question is what the program does then, and "raise and let it bubble to the user" is one answer but rarely the best one — a batch of ten thousand reviews should not stop because review 4,312 came back malformed twice.
The pattern that survives is a fallback with a marker: return a valid object of the right shape, flagged as unresolved, and keep going. The batch finishes, the flagged rows are countable, and nobody is guessing which ones failed.
What makes it work is the flag. A fallback that quietly returns {"sentiment": "neutral"} is worse than a crash, because "neutral" is a plausible answer and nobody
will ever look at it again.
Your turn: write safe_extract(question) returning (payload, ok) — the parsed
dictionary and True, or the fallback and False.
You start from this, and edit it in the browser:
import json
import fake_llm
RULE = "Reply with JSON only. No preamble."
FALLBACK = {"topic": None, "sentiment": None, "word_count": 0}
def safe_extract(question):
"""Return (payload, ok). Never raises."""
return (FALLBACK, False)