Embeddings and Vector Search · Vectors · lesson 3 of 8
Ranking by similarity
about 16 minutes · free · runs in your browser
Search is a sort
Semantic search, stripped of the marketing, is three lines: embed the query, score every document, return the best few. A vector database adds an index so that "score every document" stays fast at a million rows — but the arithmetic is the same arithmetic.
scored = [(cosine(query_vec, doc_vec), text) for text, doc_vec in store]
scored.sort(reverse=True)
One detail worth stealing: sort by score descending and keep the score alongside the text. Code that throws the scores away cannot tell "the best of a good set" from "the best of a bad set", and those need completely different behaviour — the second one should often answer "I do not know".
Your turn: write top_k(query, documents, k=2) returning the k closest
documents, best first. Return the text only; the next lesson keeps the scores.
You start from this, and edit it in the browser:
import fake_embeddings
def cosine(a, b):
return sum(x * y for x, y in zip(a, b))
def top_k(query, documents, k=2):
"""Return the k documents closest to the query, best first."""
return []