LLM APIs and Prompting · Cost and Reliability · lesson 7 of 8
Retries and exponential backoff
about 20 minutes · free · runs in your browser
Step 1 of 2
Some failures are worth another go
An LLM API fails in two kinds of way, and telling them apart is the whole job:
| Worth retrying | Not worth retrying |
|---|---|
RateLimitError — too many requests, briefly | ContentFilterError — the model declined |
APITimeoutError — no answer in time | a 400 — your request is malformed |
a 5xx APIError — their problem | a 401 — your key is wrong |
Retrying the right-hand column turns one fast, accurate failure into four slow ones with the real cause buried at the bottom of a stack trace. Retrying the left-hand column is free reliability: the second attempt usually works.
fake_llm lets you script failures so you can write the loop against real exceptions
rather than a description of one:
fake_llm.reset()
fake_llm.fail_next(2, kind="rate_limit") # the next two calls raise
Your turn: write ask_with_retry(question, attempts=3) that retries a rate limit or
a timeout and gives up immediately on anything else, re-raising it unchanged.
You start from this, and edit it in the browser:
import fake_llm
def ask_with_retry(question, attempts=3):
"""Ask, retrying only the failures another attempt could fix."""
return fake_llm.chat([{"role": "user", "content": question}])[
"choices"
][0]["message"]["content"]
Step 2 of 2
Backing off, and listening when told how long
Retrying immediately is barely a retry. Every client hitting the same limit at the same moment retries at the same moment, and the second wave is as doomed as the first — the thundering herd, and it is how a brief limit becomes an outage.
Exponential backoff spreads them: wait 1, then 2, then 4. And when the error carries a
retry_after, use it — the server has told you exactly how long, and guessing when you
have been told is not resilience.
except fake_llm.RateLimitError as exc:
delay = exc.retry_after or base * (2 ** attempt)
Your turn: write delays(attempts, base=1) returning the list of waits a client
would use for that many failures, doubling each time. Return the waits only — this
exercise is the arithmetic, not the sleeping, because a test that actually slept would
take as long as the backoff it is testing.
You start from this, and edit it in the browser:
def delays(attempts, base=1):
"""The wait before each retry: base, then double, and so on."""
return []