A repeatable way to score the system on a fixed set of real questions with known good answers, run on every change.
You have done this if
You kept 200 real questions with expected sources and ran them in CI before merging prompt changes.
Say it in a review
Every prompt, model or index change runs against the golden set and can't merge on a regression.
On the AI Application map Observability
Read Your golden set is a product, not a spreadsheet · If it runs after the deploy, it is a postmortem