LLM APIs and Prompting · Cost and Reliability · lesson 8 of 8
Rate limits and spending caps
about 18 minutes · free · runs in your browser
A budget that stops before the bill does
Retries handle failure. A budget handles success — specifically, the runaway loop that succeeds several thousand times overnight and is discovered on the invoice.
The shape is always the same: track spend, check before the call, refuse rather than truncate. Refusing is the important half. A budget that keeps calling and warns in a log is a log entry, not a budget.
Your turn: write a Budget class holding a cap in cents.
budget.ask(question)makes a call, adds its cost, and returns the text.- Once
spent_centswould exceed the cap,askraisesBudgetExceededwithout making the call — the point is to not spend the money, not to notice afterwards. budget.spent_centsreports the total so far.
Price the calls with the same per-million arithmetic as before; pylearn-small is the
model.
You start from this, and edit it in the browser:
import fake_llm
PER_MILLION = 1_000_000
MODEL = "pylearn-small"
class BudgetExceeded(Exception):
"""Raised when a call would take the total past the cap."""
class Budget:
def __init__(self, cap_cents):
self.cap_cents = cap_cents
self.spent_cents = 0
def ask(self, question):
response = fake_llm.chat([{"role": "user", "content": question}])
return response["choices"][0]["message"]["content"]