Six weeks after launch, someone forwards a screenshot. The assistant quoted a discount policy that was replaced in March. Ask how long it had been doing that and nobody can answer, because the only quality signal in the system is a person getting annoyed enough to escalate.
That is the default state of most shipped RAG. It launches with a demo, and the demo turns out to be the last evaluation it ever gets.
The fix starts with two scores instead of one. Retrieval is graded on whether the right documents came back. Generation is graded on whether the answer is supported by the documents that did come back. Collapse them into a single "was this answer good" rating and you learn that something is wrong while learning nothing about where.
Groundedness is the generation-side metric worth building first. Decompose an answer into claims and check each claim against the retrieved context. A judge model does this well enough to run continuously, and it catches the specific failure that destroys trust: a correct-sounding sentence that no retrieved document actually supports.
Citations are a per-claim contract
"Sources: [1][2][3]" at the bottom of an answer is decoration.
A citation is a promise that a specific claim came from a specific place, and it holds only when three things are true. The cited chunk contains the claim. The link opens for the user who received the answer, under their permissions rather than the service account's. The anchor lands on the relevant section rather than page one of a 90-page PDF.
Sample live answers and check all three by hand at first. The pass rate is usually lower than anyone expects, and the failures cluster in exactly the answers where the model was synthesising across several sources.
Operations is the part that runs after launch
A RAG dashboard carries more than latency:
- Groundedness and citation-validity rates over a sampled slice of live traffic
- Retrieval recall against the golden set, re-run on every index or embedding change
- Freshness: age of the oldest stale document, and deletion propagation time
- Cost per answer alongside p95 latency, because context grows quietly and both move together
- Thumbs-down volume routed to a review queue a named person actually reads
Then there is the playbook. When an answer is wrong, someone should be able to say within the hour whether the cause was a missing document, a stale one, a retrieval miss or a generation failure, and pull the offending source out of the index while it gets fixed. Without a written path, every bad answer becomes an improvised investigation run by whoever is free.
Closing thought
You do not ship RAG once. It is a data product that degrades on its own schedule, and instrumentation is the only thing that tells you it has started.
More on these topics
Checklist · · 12 checks
RAG production-readiness checklist
Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.
Checklist · · 24 checks
Working with Claude, practices that hold up
Prompting, agents and tools, Claude Code, evaluation and safety. The habits that make Claude-based systems reliable, as a checklist you can run against your own setup.
Explainer · · 2 min read
Real users ask questions your test set never imagined
Offline evals cover the questions you thought of. Production tells you the ones you did not.
Discussion