RAG Systems · Answer Quality · lesson 5 of 8
Budgeting the context window
about 20 minutes · free · runs in your browser
What fits, and what gets cut
Retrieval gives you ten candidate chunks. The context window holds four of them. Something has to be dropped, and how you choose is the difference between a good answer and a confident wrong one.
Take chunks in rank order until the budget is spent, and stop — do not truncate the chunk that does not fit. Half a passage is worse than no passage: it reads as complete, so the model treats a sentence fragment as the whole policy.
budget = 500 # tokens available for context
used = 0
for chunk in ranked:
cost = count_tokens(chunk)
if used + cost > budget:
break # not "truncate and continue"
used += cost
Leave room for the answer too. The window is shared between prompt and completion, and a prompt that fills it entirely leaves the model no space to reply — which arrives as a truncated answer rather than an error, and is diagnosed as a bad model.
Your turn: write fit_context(chunks, budget) returning the chunks that fit, in
order, plus the tokens used.
You start from this, and edit it in the browser:
import fake_llm
def fit_context(chunks, budget):
"""Return (kept_chunks, tokens_used) for the chunks that fit in budget."""
return ([], 0)