LLM APIs and Prompting · Cost and Reliability · lesson 5 of 8
Counting tokens before you send
about 14 minutes · free · runs in your browser
Step 1 of 2
The unit you are billed in
A token is roughly four characters of English — not a word, not a character. "Hello" is one token; "antidisestablishmentarianism" is several. Every provider counts slightly differently, so the arithmetic matters more than the constant.
Two reasons to count before sending rather than reading usage after:
- Context limits. Every model has a ceiling, and going over it is a hard error, not a
truncation.
fake_llm.MODELScarries the ceiling for each model. - Cost control. Knowing a request will cost 40,000 tokens is only useful before you pay for it.
fake_llm.count_tokens("Hello there") # about 3
fake_llm.MODELS["pylearn-small"]["context"] # the ceiling
Your turn: write fits(text, model) returning True when the text is inside that
model's context window. Look the ceiling up rather than hard-coding it — a model's limit
changes with a version bump, and a number typed into your code does not.
You start from this, and edit it in the browser:
import fake_llm
def fits(text, model):
"""True when text is within the model's context window."""
return True
Step 2 of 2
A cost meter
Money is integer cents, never floats. 0.1 + 0.2 is not 0.3, and a rounding error
that lands in a bill is a support ticket you cannot argue with.
fake_llm.MODELS prices each model in cents per million tokens, input and output
separately, because output is the dearer half everywhere:
fake_llm.MODELS["pylearn-small"]
# {"prompt_cents_per_m": 25, "completion_cents_per_m": 75, "context": 8000}
Per-million pricing is what providers publish, and it is why cost code multiplies before it divides: divide first and every small request rounds to zero.
Your turn: write cost_cents(usage, model) returning the whole-cent cost of one
call, rounded up. Rounding up is the honest direction for a meter — under-reporting your
own spend is how a budget alarm never fires.
You start from this, and edit it in the browser:
import fake_llm
def cost_cents(usage, model):
"""Whole cents for one call, rounded up. usage is the response's usage dict."""
return 0