RAG Systems · Ingestion · lesson 3 of 8
Metadata and filtering
about 16 minutes · free · runs in your browser
Similarity is not permission
A chunk is not just text. It came from a document, which has an owner, a date, a version and often an audience — and the moment your corpus contains more than one customer's data, retrieving by similarity alone is a data leak waiting for the right question.
{"text": "...", "source": "handbook.pdf", "tenant": "acme", "updated": "2026-08-01"}
Filter first, then rank. Not the other way round: taking the top 5 by similarity and then discarding the ones the user may not see gives you between zero and five results, varying by question, which reads as a flaky search rather than a security boundary.
Your turn: write search(query_vector, chunks, tenant, k=2) that keeps only the
chunks belonging to tenant, ranks those by cosine similarity, and returns the top
k texts.
You start from this, and edit it in the browser:
import fake_embeddings
def cosine(a, b):
return sum(x * y for x, y in zip(a, b))
def search(query_vector, chunks, tenant, k=2):
"""Filter to the tenant's chunks, then rank what remains."""
return []