Embeddings and Vector Search
Embeddings
import fake_embeddings
v = fake_embeddings.embed("bake the dough in a hot oven") # a list of DIM floats
fake_embeddings.embed_many(texts) # one call per text, billed per text
- Same text, same vector — it is a function, not a sample. That is what makes a plain dictionary a correct cache with no invalidation to get wrong.
- Fixed width, whatever the input length, so any two texts are comparable.
- Query and documents must come from the same model, or the vectors live in unrelated spaces.
Cosine similarity
import numpy as np
a, b = np.array(v1), np.array(v2)
similarity = float(a @ b / (np.linalg.norm(a) * np.linalg.norm(b)))
| Value | Meaning |
|---|---|
| 1 | same direction — same meaning |
| 0 | unrelated |
| −1 | opposite |
Both vectors unit length? The denominator is 1, so cosine is the dot product.
Search is a sort
scored = [(cosine(q, vec), text) for text, vec in store]
scored.sort(key=lambda pair: pair[0], reverse=True)
return scored[:k]
Keep the scores. Without them you cannot tell the best of a good set from the best of a bad one — and the second should often answer "I do not know".
Embed the corpus once, at index time. Re-embedding every document per query is the default mistake and it is invisible until the bill.
Chunking
| Too small | Too large |
|---|---|
| the answer splits across two chunks | the vector averages several topics |
| no context to interpret the text | the model reads a page to use a sentence |
step = max(1, size - overlap) # guard the step, or the loop never ends
Overlap repeats the tail of each chunk at the head of the next, so a sentence spanning a boundary survives whole in at least one piece.
BM25
idf = math.log(1 + (total - matching + 0.5) / (matching + 0.5))
score += idf * (tf * (k1 + 1)) / (tf + k1 * (1 - b + b * dl / avgdl))
- idf — rare terms carry the signal; "the" scores near zero with no stop-word list.
- k1 = 1.5 — saturation: the twentieth mention adds almost nothing.
- b = 0.75 — length normalisation, so long documents do not win by accident.
What each method cannot do
| Query | Semantic | Keyword |
|---|---|---|
INV-2024-0142 | no signal at all | exact hit |
| "car insurance" vs "vehicle cover" | finds it | scores zero |
This is the whole argument for hybrid retrieval: not more recall in general, but covering a known, specific blind spot each way.
Reciprocal rank fusion
for ranking in rankings:
for position, doc_id in enumerate(ranking):
scores[doc_id] = scores.get(doc_id, 0.0) + 1.0 / (k + position + 1)
Ranks, not scores — cosine and BM25 are not on the same scale and normalising them
reweights the mix on every query. k = 60 is the standard constant.
Third in both rankings (2/63) beats first in one (1/61): agreement is stronger evidence than enthusiasm.
Drop zero-scoring documents from the keyword ranking before fusing, or they collect credit for appearing in a list they do not belong in.