Embeddings and Vector Search · Retrieval · lesson 5 of 8
Chunking a document
about 18 minutes · free · runs in your browser
The unit of retrieval is not the document
Embedding a whole manual gives you one vector for forty pages, and it points at the average of everything in them — which is to say, at nothing. Retrieval works on chunks, and choosing the chunk is the single highest-leverage decision in a retrieval system.
The trade-off never goes away:
| Chunks too small | Chunks too large |
|---|---|
| the answer is split across two of them | the vector averages several topics |
| each one lacks the context to be understood | the model reads a page to use a sentence |
Overlap is the standard mitigation. Repeat the last few words of each chunk at the start of the next, so a sentence spanning a boundary survives in at least one piece. It costs storage, which is cheap, and prevents the failure where the right answer exists in the corpus and is retrievable from neither half.
Your turn: write chunk(words, size, overlap) taking a list of words and returning
lists of at most size words, each starting size - overlap words after the last.
You start from this, and edit it in the browser:
def chunk(words, size, overlap=0):
"""Split words into overlapping chunks of at most `size` words."""
return []