RAG Systems · Answer Quality · lesson 4 of 8
Measuring retrieval with recall@k
about 18 minutes · free · runs in your browser
If the right chunk is not retrieved, nothing else matters
When a RAG answer is wrong there are two candidates: retrieval brought back the wrong passages, or the model misused the right ones. They have completely different fixes, and guessing between them is how weeks disappear.
Recall@k settles it. Take a set of questions whose answering chunk you know, and
measure how often that chunk is in the top k. If recall@5 is 0.4, no amount of prompt
engineering will save you — the evidence simply is not reaching the model 60% of the time.
It is the cheapest measurement in this course: no model call, no judgement, just set membership. Measure it before touching the prompt.
Your turn: write recall_at_k(results, expected, k). results is a list of
ranked id lists, expected the id that should have been found for each. Return the
fraction found within the first k, as a float.
You start from this, and edit it in the browser:
def recall_at_k(results, expected, k):
"""Fraction of questions whose expected id appears in the top k."""
return 0.0