Embeddings and Vector Search · Retrieval · lesson 7 of 8
Where each method fails
about 16 minutes · free · runs in your browser
Neither one is the good one
The two methods fail in opposite directions, and the failures are predictable enough to plan around.
Semantic search cannot find an identifier. Ask for invoice INV-2024-0142 and the
embedding has no notion of it — the string carries no topic, so the vector is noise and the
ranking is arbitrary. Exact tokens are precisely what a keyword index is for.
Keyword search cannot find a synonym. Ask about "car insurance" of a corpus that says "vehicle cover" and BM25 scores zero: no shared terms, no match, however obviously related the two are.
This is why serious retrieval systems run both. Not for extra recall in general — for covering each other's specific, known blind spot.
Your turn: write keyword_hits(query_terms, documents) returning the indexes of the
documents containing all the query terms. This is the exact-match half, and its
strictness is the point.
You start from this, and edit it in the browser:
def keyword_hits(query_terms, documents):
"""Indexes of documents containing every query term."""
return []