LLM APIs and Prompting · Calling a Model · lesson 4 of 8
Temperature, length and stopping
about 18 minutes · free · runs in your browser
Step 1 of 2
Temperature is a dial for variety, not for quality
Temperature controls how much the model is willing to deviate from its most likely next
word. At 0 you get the same answer for the same input, every time. Raise it and the
wording moves around.
The rule of thumb that survives contact with production: anything your code parses runs at temperature 0. Classification, extraction, routing, anything with a schema. Save the variety for prose a person reads, where two phrasings of the same idea are both fine.
cold = fake_llm.chat(msg, temperature=0) # the same answer every time
warm = fake_llm.chat(msg, temperature=1) # a different phrasing
In fake_llm the variation is seeded rather than random, so a higher temperature gives a
different answer that is still the same on every run. A real model gives you a genuinely
different one — which is exactly why you cannot write a test against it, and why the code
around it has to be written so that the difference does not matter.
Your turn: write differs(question) returning True when temperature 1 produces a
different string from temperature 0 for the same question.
You start from this, and edit it in the browser:
import fake_llm
def differs(question):
"""True when temperature 1 answers differently from temperature 0."""
return False
Step 2 of 2
max_tokens is a budget, and it cuts mid-word
max_tokens caps the output. When the answer hits the cap it stops where it is —
mid-sentence, mid-word, mid-JSON — and the response comes back with HTTP 200 and a full
bill.
The only thing that tells you is finish_reason:
| Value | Meaning |
|---|---|
"stop" | the model finished, or hit a stop sequence |
"length" | it ran out of budget — the answer is incomplete |
This is the failure people ship. A truncated reply looks like a short reply, and short replies are normal, so nothing raises. Days later a JSON parse fails on one request in a hundred, and the cause is a cap set months earlier.
stop sequences are the other end of the same idea: give the model a string that means
"stop here", and the text is cut at it, with finish_reason still "stop" because
stopping was the plan.
Your turn: write complete(question, budget) returning (text, was_truncated).
You start from this, and edit it in the browser:
import fake_llm
def complete(question, budget):
"""Return (text, was_truncated) for one question under a token budget."""
return ("", False)