LLM APIs and Prompting · Calling a Model · lesson 1 of 8
What an LLM API is
about 14 minutes · free · runs in your browser
Step 1 of 2
An API that answers in prose
An LLM API is an ordinary HTTP endpoint. You post a conversation, you get a message back, and you are billed for the text that went in and the text that came out. Everything mysterious about the model is on the other side of that boundary; on this side it is a function call with a dictionary for an answer.
In this course you will use fake_llm, which has the shape of a real completions API
and answers deterministically. Python in the browser has no sockets, so there is nothing
to reach — but writing an LLM client against a deterministic stub is how you would test
one anywhere, because you cannot write an assertion against a sampler.
import fake_llm
response = fake_llm.chat([{"role": "user", "content": "What is rag?"}])
response["choices"][0]["message"]["content"] # the answer, as text
response["choices"][0]["finish_reason"] # why it stopped: "stop", "length", ...
response["usage"]["total_tokens"] # what it cost you
That nesting is not decoration. choices is a list because an API can return several
candidate answers, and finish_reason is separate from the text because how the model
stopped is a different question from what it said — a truncated answer arrives with a
perfectly good HTTP 200.
Your turn: ask the model what a token is, and pull out the reply and the reason it stopped.
You start from this, and edit it in the browser:
import fake_llm
# Ask the model: "What is a token?"
# Set reply to the assistant's text and finish to the finish_reason.
Step 2 of 2
The bill arrives with the answer
Every response carries usage, and it is split in two: the tokens you sent and the
tokens the model produced. They are priced differently — output is several times dearer
than input at every provider — so a program that tracks one number is tracking the wrong
one.
usage = response["usage"]
usage["prompt_tokens"] # what you sent
usage["completion_tokens"] # what came back
usage["total_tokens"] # the sum, for convenience
The trap: people measure cost by counting their prompt and forget that the conversation grows. Send a chat history back on every turn and the prompt tokens climb with each message, so the tenth question in a conversation can cost ten times the first — for the same question.
Your turn: write ask(question) returning a (text, prompt_tokens, completion_tokens) tuple.
You start from this, and edit it in the browser:
import fake_llm
def ask(question):
"""Return (text, prompt_tokens, completion_tokens) for one question."""
return ("", 0, 0)