Five-day tracks
5 Days of AI Evaluation
Evaluation that can block a release: what an eval may stop, golden sets maintained like products, calibrated judges, regression gates and evaluation after launch.
5 of 5 published
- 01
Part 1 · Explainer · 2 min read
An eval that cannot block a release is just a report
Decide which decision the number is allowed to stop, then build the harness around that.
On the map: Observability - 02
Part 2 · Explainer · 2 min read
Your golden set is a product, not a spreadsheet
Labels, edge cases, versions, an owner. Skip those and the number moves without anyone knowing why.
On the map: Observability - 03
Part 3 · Explainer · 2 min read
A judge that cannot name the failure is not a judge
Deterministic checks first, one rubric per criterion, and a calibration loop against human review.
On the map: Observability - 04
Part 4 · Explainer · 2 min read
If it runs after the deploy, it is a postmortem
Regression gates only work when they sit in the release path, with a threshold, a slice and an owner.
On the map: Observability - 05
Part 5 · Explainer · 2 min read
Launch day is when evaluation starts
Offline scores expire on contact with real users. Sampling, groundedness, drift and outcomes are the parts that keep paying.
On the map: Observability