RAG Systems
The loop
ingest: parse → clean → chunk → embed → store (with metadata)
answer: retrieve → filter → rank → budget → assemble → generate → verify
One rule underneath all of it: the model may only use what you put in front of it.
Ingestion
paragraphs = [" ".join(block.split()) for block in text.split("\n\n")]
paragraphs = [p for p in paragraphs if p and not p.isdigit()]
- Collapse whitespace; drop page furniture; keep paragraph boundaries.
- Chunk on paragraphs first, word counts only as a fallback.
- Store metadata beside the text:
source,tenant,updated. Text alone cannot answer "may this user see it" or "is this current".
Filtering
Filter, then rank. Ranking first and filtering after returns a different number of results per question — a flaky search and a leak in the same mistake.
Measuring
recall_at_k = found_in_top_k / total_questions
The first measurement, always. It separates "retrieval missed it" from "the model misused it", needs no model call, and stops weeks being spent on prompt engineering that cannot help.
Budgeting the window
for chunk in ranked:
if used + count_tokens(chunk) > budget:
break # stop; never truncate
Half a passage reads as a whole one. Leave headroom for the answer as well — prompt and completion share the window.
Citations
numbered = "\n".join(f"[{i}] {t}" for i, t in enumerate(passages, 1))
cited = [n for n in re.findall(r"\[(\d+)\]", answer) if 1 <= int(n) <= len(passages)]
Number the passages, require citation by number, and validate what comes back. A citation pointing past the end of the context is worse than none: it looks checkable.
An answer with no citation at all came from the model's memory, not your documents.
The four silent failures
| Failure | Symptom | Mitigation |
|---|---|---|
| chunk boundary | a confident half-answer | overlap; chunk on structure |
| stale context | last year's policy, stated firmly | store updated; re-index |
| lost in the middle | the right chunk retrieved and ignored | strongest chunks at the edges |
| conflicting sources | one of two contradictory answers | detect disagreement; surface it |
front, back = chunks[0::2], chunks[1::2][::-1]
context = front + back # best first, second-best last, weakest buried
Refusing
Two gates, both needed:
- Before the call — best similarity below a floor means do not ask. Retrieval always returns something; the top of a bad set is still the top.
- After the call — no citation means ungrounded, whatever the answer says.
if not scored or scored[0][0] < floor:
return (REFUSAL, False) # costs nothing; a fluent guess costs more than money
The most valuable answer a grounded system gives is sometimes "I do not have anything in these documents that answers that".