Evals and Shipping · Shipping · lesson 5 of 7
Caching, exact and semantic
about 18 minutes · free · runs in your browser
The cheapest call is the one you do not make
Two kinds of cache, with very different risk profiles.
Exact caching keys on the request — model, temperature, and the full messages. At temperature 0 it is free correctness: the same input would have produced the same output. The only subtlety is including everything that affects the answer in the key. Leave the model out and a cheap model's answer is served for an expensive model's request.
Semantic caching goes further: if a new question is close enough to a cached one, serve the old answer. That is a genuine trade — "how do I cancel" and "how do I not cancel" are extremely close in embedding space and want opposite answers. Semantic caching belongs behind a high threshold and, ideally, human review of what it merged.
Your turn: build Cache with key(model, temperature, messages) and
get_or_call(model, temperature, messages, call), where call is a zero-argument
function producing the answer. Everything that changes the answer must be in the key.
You start from this, and edit it in the browser:
import json
class Cache:
def __init__(self):
self.store = {}
self.hits = 0
self.misses = 0
def key(self, model, temperature, messages):
"""A key covering everything that changes the answer."""
return ""
def get_or_call(self, model, temperature, messages, call):
"""Return the cached answer, or call and store it."""
return call()