Evals and Shipping · Measuring · lesson 4 of 7
Prompt regression tests in CI
about 20 minutes · free · runs in your browser
A number, a baseline, and a threshold
An eval you run by hand when you remember is a demo. The version that changes behaviour runs on every prompt change and fails the build when the score drops.
Which means it needs a baseline — the score the current prompt achieves, committed to the repository like any other expectation — and a threshold, because a suite that must never regress by a single case will be disabled within a month.
baseline: 0.92 the score today, committed
tolerance: 0.02 natural variation, agreed in advance
fail if: score < baseline - tolerance
The tolerance is the part people get wrong in both directions. Zero tolerance turns one flaky case into a blocked deploy and the suite gets skipped. A large tolerance lets a real regression through unnoticed. Pick it from the variation you actually observe across runs.
Your turn: write gate(score, baseline, tolerance=0.02) returning ("pass", None)
or ("fail", message). Improving on the baseline is always a pass, and the message must
name both numbers — a failing build that does not say by how much sends someone digging
through logs.
You start from this, and edit it in the browser:
def gate(score, baseline, tolerance=0.02):
"""Return ('pass', None) or ('fail', message)."""
return ("pass", None)