Skip to content

Structured Output and Tool Calling

The boundary, in four lines

response = fake_llm.chat(messages, temperature=0)     # 1. ask, with the shape stated
payload = json.loads(response["choices"][0]["message"]["content"])   # 2. parse
check(payload, SCHEMA)                                # 3. validate
return payload                                        # 4. only now is it data

Everything before the return is the boundary. Everything after may assume the data is what it claims to be — which is the whole reason the boundary exists.

Asking for a shape

ContractInstructionCheck
one label"Reply with exactly one word: positive, negative or neutral."membership in a set
JSON"Reply with JSON only. No preamble, no code fence."json.loads
a schemalist the fields and types in the promptfield by field

Stating the shape is half the job. The other half is never trusting that it was followed.

Validating

SCHEMA = {"topic": str, "sentiment": str, "word_count": int}

for field, expected in SCHEMA.items():
    if field not in payload:
        raise ValueError("missing field: " + field)
    if not isinstance(payload[field], expected):
        raise ValueError("wrong type for " + field)
  • Name the field in the error. "Validation failed" is an investigation; "missing field: word_count" is an answer.
  • bool is a subclass of int. isinstance(True, int) is True, so a boolean slips into a numeric field unless you check for it first.
  • In production this is a Pydantic model. What it does is exactly the loop above.

Repairing

try:
    return json.loads(content)
except ValueError as exc:
    messages += [
        {"role": "assistant", "content": content},
        {"role": "user", "content": "That was not valid JSON: " + str(exc) + ". JSON only."},
    ]
    # ...and call again

Re-ask, do not patch. A regex that strips preambles is a guess about a shape you do not control. When repair runs out of attempts, fall back to a flagged object of the right shape — never to a plausible value nobody will re-examine.

Tools

tools = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Look up the current weather for a city",   # this is the prompt
        "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"],
        },
    },
}]

The round trip:

response = fake_llm.chat(messages, tools=tools)
message = response["choices"][0]["message"]        # content is None on a tool call

call = message["tool_calls"][0]
name = call["function"]["name"]
args = json.loads(call["function"]["arguments"])   # a JSON *string*, always
result = registry[name](**args)                    # your code runs it, not the model

messages += [message, {"role": "tool", "content": str(result)}]
final = fake_llm.chat(messages)                    # no tools — it has what it needed

The three ways a tool call arrives broken

SymptomCauseHandle it by
message["tool_calls"] missingno tool fitted; the model answered in prosereturning the prose
json.loads raisestruncated by max_tokensretrying, never guessing
registry[name] KeyErrorthe model named a tool that does not existchecking membership first
fake_llm.break_next("truncated_tool_call")   # to test the second

Habits

  • Temperature 0 for everything in this course. All of it is parsed.
  • Validate at the edge, once.
  • Re-ask to repair; fall back with a flag; never invent a value.
  • Check the tool name against your registry before calling it.
  • Drop the tools from the final call, or the round trip has no stopping condition.

This cheatsheet is the summary. If you want to build it yourself, the Structured Output and Tool Calling course walks you through it in the browser — the first lesson is free.