Evals and Shipping · Shipping · lesson 6 of 7
Latency and model routing
about 20 minutes · free · runs in your browser
Spend the big model where it earns its price
Most requests do not need the expensive model. Classification, extraction from a clean document, a short factual lookup — a small model handles them at a fraction of the price and a fraction of the wait. Routing is the practice of noticing that.
A router does not need to be clever. Cheap, legible rules beat a learned classifier that nobody can explain when it misroutes:
- short input, simple task → small model
- long context, or a task the small model has measurably failed → large model
- escalate on failure: try small, validate, retry on the large one if validation fails
Escalation is the pattern worth internalising, because it makes routing safe. The small model handling 90% of traffic and escalating the rest costs less than the large model handling everything, and is more accurate than the small model alone.
Your turn: write route(text, task, failed_before=False) returning
"pylearn-small" or "pylearn-large". Use the large model when the text is over 400
tokens, when the task is "reasoning", or when a previous attempt failed.
You start from this, and edit it in the browser:
import fake_llm
SMALL = "pylearn-small"
LARGE = "pylearn-large"
LONG_INPUT_TOKENS = 400
HARD_TASKS = {"reasoning"}
def route(text, task, failed_before=False):
"""Choose a model for this request."""
return SMALL