Turn meaning into geometry, then find the right passage twice over.
8 lessons · about 7 hours · free
An embedding puts a text somewhere in space so that distance means similarity. This course builds the whole retrieval stack by hand: cosine similarity in NumPy, nearest-neighbour ranking, batching and caching the embedding calls you are billed for, chunking a document without slicing sentences in half, BM25 from scratch, and reciprocal rank fusion to combine the two. It spends as much time on where each method fails — semantic search cannot find an invoice number, keyword search cannot find a synonym — because knowing that is what makes hybrid retrieval an obvious choice rather than a fashionable one.
What an embedding is geometrically, how to compare two, and what the calls cost.
Chunking, BM25, the failure modes of each method, and combining them properly.