Skip to content

LLM APIs and Prompting

In PyLearn the import is fake_llm, because Python in the browser has no sockets. Everything below is the shape you would write against a real completions API.

One call

import fake_llm

response = fake_llm.chat(
    [
        {"role": "system", "content": "Answer in one sentence."},
        {"role": "user", "content": "What is rag?"},
    ],
    model="pylearn-small",   # fake_llm.MODELS lists them with prices and context limits
    temperature=0,           # 0 for anything you parse
    max_tokens=200,          # caps the answer; exceeding it sets finish_reason "length"
    stop=None,               # a string or list that ends the answer early
)

text = response["choices"][0]["message"]["content"]
why = response["choices"][0]["finish_reason"]     # "stop" | "length" | "tool_calls"
usage = response["usage"]                         # prompt_tokens, completion_tokens, total

The messages list

RoleFor
systemstanding rules — role, format, prohibitions. Sent every call.
userwhat the person said
assistantwhat the model said last turn

The API keeps no state. The conversation is a list you own and re-send, so the prompt grows every turn — and so does its cost.

messages.append({"role": "user", "content": question})
reply = fake_llm.chat(messages)["choices"][0]["message"]["content"]
messages.append({"role": "assistant", "content": reply})

A Client holds the defaults so they are not retyped at every call site:

client = fake_llm.Client(system="Answer in one word.", temperature=0)
client.ask("Is this positive or negative?")   # the text, straight back

Sampling

SettingEffect
temperature=0repeatable. Use it for anything parsed.
temperature>0varies the wording. Prose only.
max_tokenscaps output. Watch finish_reason.
stopcuts the text at a marker you choose

Money

Prices are cents per million tokens, input and output separately:

price = fake_llm.MODELS["pylearn-small"]
total = (
    usage["prompt_tokens"] * price["prompt_cents_per_m"]
    + usage["completion_tokens"] * price["completion_cents_per_m"]
)
cents = -(-total // 1_000_000)   # multiply first, divide last, round up

Divide per token and every small call rounds to zero. Never use floats for money.

Streaming

for chunk in fake_llm.chat(messages, stream=True):
    choice = chunk["choices"][0]
    piece = choice["delta"].get("content", "")   # .get — the last chunk has none
    if choice["finish_reason"]:
        done = choice["finish_reason"]

The final chunk carries finish_reason and usage and no content. Indexing delta["content"] raises on it, every single time.

Failing well

RetryDo not retry
RateLimitError (429)ContentFilterError — the model declined
APITimeoutError400 — the request is wrong
APIError 5xx401 — the key is wrong
try:
    response = fake_llm.chat(messages)
except fake_llm.RateLimitError as exc:
    delay = exc.retry_after or base * (2 ** attempt)   # obey the server if it said

Scripting failures, so a retry loop can be tested against real exceptions:

fake_llm.reset()
fake_llm.fail_next(2, kind="rate_limit")   # or "timeout", "server", "filtered"
fake_llm.break_next("malformed_json")      # succeeds, bills you, will not parse
fake_llm.call_count()                      # what your loop actually did

State resets between runs, not between the tests inside one run — call reset() first in any test that counts calls.

The habits worth keeping

  • Temperature 0 for anything with a schema.
  • Read finish_reason before trusting the text.
  • Count tokens before sending, not after paying.
  • Check the budget before the call; a cap that notices afterwards is a receipt.
  • Log the exact messages sent — that is where the prompting bug is.

This cheatsheet is the summary. If you want to build it yourself, the LLM APIs and Prompting course walks you through it in the browser — the first lesson is free.