New here? Start this series at Part 1: Logging the answer tells you almost nothing
- Part 1: Logging the answer tells you almost nothing
- Part 2: Your AI feature has unit economics whether you measured them or not
- Part 3: Real users ask questions your test set never imagined
- Part 4: You cannot roll back a prompt you never versioned
- Part 5: Nothing broke, and the system is still getting worse
Your golden set has two hundred questions in it, and every one was written by somebody on the team. Which means every one is a question asked by a person who already understands the product.
Real users ask why this is different from last month. They paste half a spreadsheet. They ask a follow-up that only makes sense given the previous three turns. None of that is in the golden set, and it will keep passing while those cases fail.
Sample production before it surprises you
Pick a rate and review real traffic on a schedule. Small is fine. Fifty conversations a week read properly beats five thousand skimmed.
Privacy is the constraint that stops most teams doing this at all, so design around it instead of abandoning it. Review inside a controlled tool rather than exporting to a spreadsheet. Redact identifiers on the way in. Then stratify, so you are not just reading the median: over-sample long conversations, high-cost runs, sessions with retries, anything that tripped a guardrail. The boring middle teaches you nothing.
Three signals users hand you for free
Explicit feedback. Thumbs are weak but not useless. The value sits almost entirely in the optional comment box and in the ratio moving rather than the absolute number.
Corrections. When a user edits the output before using it, that edit is a labelled pair: what the system produced, what a person actually wanted. It is the highest quality signal in the product, and most teams discard it because the edit happens in a different component from the generation.
Support labels. Every escalation carries a reason. Free text gives you anecdotes. A small controlled vocabulary matching your failure taxonomy (wrong retrieval, stale data, refused incorrectly, fabricated detail, tone) gives you a distribution you can rank work against.
Attach the trace ID to all three, or the signal arrives without its evidence.
Closing the loop is somebody's job
Collecting feedback that never becomes test data is theatre. Someone has to own the path from signal to dataset: triage the week's flagged cases, decide which represent a class rather than a one-off, write the expected behaviour, add them to the regression set.
That last step is where it stalls, because writing the expected output is genuine work and nobody's sprint has room for it. Budget the hours or accept that your eval set stays frozen at launch quality.
Aggregates hide the regression
A release that improves the mean can still break one tenant, one language, one document type, one risk tier. Slice by workflow, tenant, model version and risk level, and set your rollback threshold on the worst slice rather than the average.
Decide those thresholds while calm. What drop in slice quality reverses a rollout? What rate of safety flags halts expansion to the next ring? Written down in advance, they are policy. Debated mid-rollout, they are whatever the loudest person in the channel believes.
Which slice of your traffic could regress by a third before anyone noticed?
More on these topics
Deep dive · · 6 min read
Every question your AI readiness review asks was answered months ago
The last six days of a sixty-day series, and the pattern is that operability gets bought early or it does not get bought at all.
Checklist · · 12 checks
RAG production-readiness checklist
Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.
Checklist · · 24 checks
Working with Claude, practices that hold up
Prompting, agents and tools, Claude Code, evaluation and safety. The habits that make Claude-based systems reliable, as a checklist you can run against your own setup.
Discussion