LLM APIs and Prompting · Cost and Reliability · lesson 6 of 8
Streaming an answer
about 16 minutes · free · runs in your browser
Words as they arrive
A model that takes eight seconds to answer feels broken. The same model streaming its first words in three hundred milliseconds feels fast, and the total time is identical — the difference is entirely in when the first character appears.
With stream=True the call returns a generator of chunks rather than one response.
Each chunk carries a delta: the piece of text produced since the last one.
for chunk in fake_llm.chat(messages, stream=True):
delta = chunk["choices"][0]["delta"]
piece = delta.get("content", "")
The trap is the last chunk. It carries finish_reason and usage and no
content at all — because it is telling you the answer ended, not adding to it. Code
written as chunk["choices"][0]["delta"]["content"] raises KeyError on the final
chunk of every successful stream, which is the most reliable way to ship something that
works in the demo and breaks on every real request.
Your turn: write collect(question) returning (text, finish_reason) by reading
the stream to the end.
You start from this, and edit it in the browser:
import fake_llm
def collect(question):
"""Read a streamed answer and return (text, finish_reason)."""
return ("", "")