Embeddings and Vector Search · Vectors · lesson 4 of 8
Batching and caching what you pay for
about 16 minutes · free · runs in your browser
Embeddings are cheap and you will still overspend
Embedding calls are billed per token like everything else. The naive retrieval loop re-embeds the entire corpus on every single query — a thousand documents, a thousand calls, for one question — and it is fast enough in development that nobody notices until the bill.
Two fixes, in order of importance:
- Embed the corpus once, up front. Documents do not change between queries. Store the vectors beside the text.
- Cache by exact text. Embedding is a pure function, so the same string always gives the same vector — a dictionary keyed on the text is a correct cache, with no expiry to reason about.
_cache = {}
def cached_embed(text):
if text not in _cache:
_cache[text] = fake_embeddings.embed(text)
return _cache[text]
Your turn: build Index — embed the documents once in __init__, and answer
search(query, k) without re-embedding any of them.
You start from this, and edit it in the browser:
import fake_embeddings
class Index:
def __init__(self, documents):
self.documents = list(documents)
def search(self, query, k=2):
"""Return the k best documents, best first."""
return []