LLM APIs and Prompting
In PyLearn the import is fake_llm, because Python in the browser has no sockets.
Everything below is the shape you would write against a real completions API.
One call
import fake_llm
response = fake_llm.chat(
[
{"role": "system", "content": "Answer in one sentence."},
{"role": "user", "content": "What is rag?"},
],
model="pylearn-small", # fake_llm.MODELS lists them with prices and context limits
temperature=0, # 0 for anything you parse
max_tokens=200, # caps the answer; exceeding it sets finish_reason "length"
stop=None, # a string or list that ends the answer early
)
text = response["choices"][0]["message"]["content"]
why = response["choices"][0]["finish_reason"] # "stop" | "length" | "tool_calls"
usage = response["usage"] # prompt_tokens, completion_tokens, total
The messages list
| Role | For |
|---|---|
system | standing rules — role, format, prohibitions. Sent every call. |
user | what the person said |
assistant | what the model said last turn |
The API keeps no state. The conversation is a list you own and re-send, so the prompt grows every turn — and so does its cost.
messages.append({"role": "user", "content": question})
reply = fake_llm.chat(messages)["choices"][0]["message"]["content"]
messages.append({"role": "assistant", "content": reply})
A Client holds the defaults so they are not retyped at every call site:
client = fake_llm.Client(system="Answer in one word.", temperature=0)
client.ask("Is this positive or negative?") # the text, straight back
Sampling
| Setting | Effect |
|---|---|
temperature=0 | repeatable. Use it for anything parsed. |
temperature>0 | varies the wording. Prose only. |
max_tokens | caps output. Watch finish_reason. |
stop | cuts the text at a marker you choose |
Money
Prices are cents per million tokens, input and output separately:
price = fake_llm.MODELS["pylearn-small"]
total = (
usage["prompt_tokens"] * price["prompt_cents_per_m"]
+ usage["completion_tokens"] * price["completion_cents_per_m"]
)
cents = -(-total // 1_000_000) # multiply first, divide last, round up
Divide per token and every small call rounds to zero. Never use floats for money.
Streaming
for chunk in fake_llm.chat(messages, stream=True):
choice = chunk["choices"][0]
piece = choice["delta"].get("content", "") # .get — the last chunk has none
if choice["finish_reason"]:
done = choice["finish_reason"]
The final chunk carries finish_reason and usage and no content. Indexing
delta["content"] raises on it, every single time.
Failing well
| Retry | Do not retry |
|---|---|
RateLimitError (429) | ContentFilterError — the model declined |
APITimeoutError | 400 — the request is wrong |
APIError 5xx | 401 — the key is wrong |
try:
response = fake_llm.chat(messages)
except fake_llm.RateLimitError as exc:
delay = exc.retry_after or base * (2 ** attempt) # obey the server if it said
Scripting failures, so a retry loop can be tested against real exceptions:
fake_llm.reset()
fake_llm.fail_next(2, kind="rate_limit") # or "timeout", "server", "filtered"
fake_llm.break_next("malformed_json") # succeeds, bills you, will not parse
fake_llm.call_count() # what your loop actually did
State resets between runs, not between the tests inside one run — call reset() first
in any test that counts calls.
The habits worth keeping
- Temperature 0 for anything with a schema.
- Read
finish_reasonbefore trusting the text. - Count tokens before sending, not after paying.
- Check the budget before the call; a cap that notices afterwards is a receipt.
- Log the exact messages sent — that is where the prompting bug is.