Skip to content

Embeddings and Vector Search

Embeddings

import fake_embeddings

v = fake_embeddings.embed("bake the dough in a hot oven")   # a list of DIM floats
fake_embeddings.embed_many(texts)                           # one call per text, billed per text
  • Same text, same vector — it is a function, not a sample. That is what makes a plain dictionary a correct cache with no invalidation to get wrong.
  • Fixed width, whatever the input length, so any two texts are comparable.
  • Query and documents must come from the same model, or the vectors live in unrelated spaces.

Cosine similarity

import numpy as np

a, b = np.array(v1), np.array(v2)
similarity = float(a @ b / (np.linalg.norm(a) * np.linalg.norm(b)))
ValueMeaning
1same direction — same meaning
0unrelated
−1opposite

Both vectors unit length? The denominator is 1, so cosine is the dot product.

Search is a sort

scored = [(cosine(q, vec), text) for text, vec in store]
scored.sort(key=lambda pair: pair[0], reverse=True)
return scored[:k]

Keep the scores. Without them you cannot tell the best of a good set from the best of a bad one — and the second should often answer "I do not know".

Embed the corpus once, at index time. Re-embedding every document per query is the default mistake and it is invisible until the bill.

Chunking

Too smallToo large
the answer splits across two chunksthe vector averages several topics
no context to interpret the textthe model reads a page to use a sentence
step = max(1, size - overlap)   # guard the step, or the loop never ends

Overlap repeats the tail of each chunk at the head of the next, so a sentence spanning a boundary survives whole in at least one piece.

BM25

idf = math.log(1 + (total - matching + 0.5) / (matching + 0.5))

score += idf * (tf * (k1 + 1)) / (tf + k1 * (1 - b + b * dl / avgdl))
  • idf — rare terms carry the signal; "the" scores near zero with no stop-word list.
  • k1 = 1.5 — saturation: the twentieth mention adds almost nothing.
  • b = 0.75 — length normalisation, so long documents do not win by accident.

What each method cannot do

QuerySemanticKeyword
INV-2024-0142no signal at allexact hit
"car insurance" vs "vehicle cover"finds itscores zero

This is the whole argument for hybrid retrieval: not more recall in general, but covering a known, specific blind spot each way.

Reciprocal rank fusion

for ranking in rankings:
    for position, doc_id in enumerate(ranking):
        scores[doc_id] = scores.get(doc_id, 0.0) + 1.0 / (k + position + 1)

Ranks, not scores — cosine and BM25 are not on the same scale and normalising them reweights the mix on every query. k = 60 is the standard constant.

Third in both rankings (2/63) beats first in one (1/61): agreement is stronger evidence than enthusiasm.

Drop zero-scoring documents from the keyword ranking before fusing, or they collect credit for appearing in a list they do not belong in.

This cheatsheet is the summary. If you want to build it yourself, the Embeddings and Vector Search course walks you through it in the browser — the first lesson is free.