Skip to content

RAG Systems

The loop

ingest:   parse → clean → chunk → embed → store (with metadata)
answer:   retrieve → filter → rank → budget → assemble → generate → verify

One rule underneath all of it: the model may only use what you put in front of it.

Ingestion

paragraphs = [" ".join(block.split()) for block in text.split("\n\n")]
paragraphs = [p for p in paragraphs if p and not p.isdigit()]
  • Collapse whitespace; drop page furniture; keep paragraph boundaries.
  • Chunk on paragraphs first, word counts only as a fallback.
  • Store metadata beside the text: source, tenant, updated. Text alone cannot answer "may this user see it" or "is this current".

Filtering

Filter, then rank. Ranking first and filtering after returns a different number of results per question — a flaky search and a leak in the same mistake.

Measuring

recall_at_k = found_in_top_k / total_questions

The first measurement, always. It separates "retrieval missed it" from "the model misused it", needs no model call, and stops weeks being spent on prompt engineering that cannot help.

Budgeting the window

for chunk in ranked:
    if used + count_tokens(chunk) > budget:
        break        # stop; never truncate

Half a passage reads as a whole one. Leave headroom for the answer as well — prompt and completion share the window.

Citations

numbered = "\n".join(f"[{i}] {t}" for i, t in enumerate(passages, 1))
cited = [n for n in re.findall(r"\[(\d+)\]", answer) if 1 <= int(n) <= len(passages)]

Number the passages, require citation by number, and validate what comes back. A citation pointing past the end of the context is worse than none: it looks checkable.

An answer with no citation at all came from the model's memory, not your documents.

The four silent failures

FailureSymptomMitigation
chunk boundarya confident half-answeroverlap; chunk on structure
stale contextlast year's policy, stated firmlystore updated; re-index
lost in the middlethe right chunk retrieved and ignoredstrongest chunks at the edges
conflicting sourcesone of two contradictory answersdetect disagreement; surface it
front, back = chunks[0::2], chunks[1::2][::-1]
context = front + back        # best first, second-best last, weakest buried

Refusing

Two gates, both needed:

  1. Before the call — best similarity below a floor means do not ask. Retrieval always returns something; the top of a bad set is still the top.
  2. After the call — no citation means ungrounded, whatever the answer says.
if not scored or scored[0][0] < floor:
    return (REFUSAL, False)   # costs nothing; a fluent guess costs more than money

The most valuable answer a grounded system gives is sometimes "I do not have anything in these documents that answers that".

This cheatsheet is the summary. If you want to build it yourself, the RAG Systems course walks you through it in the browser — the first lesson is free.