RAG Systems · Ingestion · lesson 1 of 8
The RAG loop
about 14 minutes · free · runs in your browser
Why retrieve at all
A model knows what was in its training data. It does not know your invoices, your runbook, or anything written last week — and asked anyway, it will produce something fluent and wrong, because fluency is what it optimises.
RAG fixes the input rather than the model: find the relevant passages first, put them in the prompt, and ask the question about them.
question ─→ retrieve ─→ assemble prompt ─→ model ─→ answer + citations
↑
your documents
The whole design follows from one rule: the model may only use what you put in front of it. Everything else in this course is a consequence — chunking decides what can be retrieved, budgeting decides what survives to the prompt, citations prove which passage an answer came from, and refusing is what happens when the passages do not support one.
Your turn: write build_prompt(question, passages) returning the messages list.
System message: the standing rule that the model answers only from the context. Then the
context and the question as a user message.
You start from this, and edit it in the browser:
RULE = "Answer using only the context provided."
def build_prompt(question, passages):
"""Return a messages list grounding the question in the passages."""
return []