RAG Systems · Ingestion · lesson 2 of 8
Parse, clean, chunk
about 18 minutes · free · runs in your browser
Garbage in, confident garbage out
Ingestion is unglamorous and decides everything downstream. A PDF converted badly gives you page numbers and running headers interleaved with sentences; those fragments get embedded, retrieved, and quoted back to a user as if they meant something.
Cleaning is boring and specific:
- collapse runs of whitespace, which survive most converters
- drop lines that are page furniture — bare numbers, repeated headers
- keep paragraph boundaries, because they are the natural chunk edge
Then chunk on paragraphs first, falling back to a word count only when a paragraph is too long. A chunk that follows the document's own structure needs less overlap and reads as something a person wrote, which matters when it is quoted back to them.
Your turn: write clean(text) returning tidy paragraphs — a list of strings, each
with its internal whitespace collapsed, page furniture removed, and empty entries dropped.
You start from this, and edit it in the browser:
def clean(text):
"""Return a list of tidy paragraphs."""
return [text]