Someone shows me a dashboard. Eighty-seven per cent on the internal benchmark, up from eighty-four last month. Good news, apparently.

I ask what happens if it comes back at seventy-nine next Tuesday. The answer is usually some version of "we would look into it".

Then what they have is a chart with a nice trend line.

Start from the decision and work backwards

Before designing any scoring, name the decision the score is meant to serve. Does this block a prompt change? A model swap? A rollout to a new customer segment? Each tolerates completely different evidence. A prompt tweak needs a fast regression check in CI. Opening a regulated workflow to a new market needs human review of adversarial cases and a named person willing to sign.

Then set the threshold before you look at results. This sounds pedantic until you have watched a team discover a four-point drop and spend an afternoon arguing about whether four points matters. It always matters less once everyone knows which release it would delay. Write the number down while it is still abstract.

Model score and business outcome are also not the same quantity, and treating them as one is how teams ship confident regressions. Higher scores that produce more support tickets are a failure. Track both, and be explicit about which one wins when they disagree.

Not every task fails the same way

Lumping every request into one accuracy number hides exactly the failures you care about. Split by user intent and by what an error costs.

Task typeWhat correct meansBest signalWhat it can block
ExtractionField matches the sourceDeterministic diff, offlineAny release
SummarisationNothing invented, key points keptRubric plus sampled human reviewPrompt and model changes
Advice and support answersUseful and safe for this userHuman review, online feedbackSegment rollout
Agentic actionsRight steps, stopped at the right pointTrajectory reviewExpanding autonomy

The three sources of signal do different jobs. Offline evals are cheap and repeatable, so they belong in the release path. Online evals catch the questions your golden set never imagined. Human review is expensive and irreplaceable wherever "wrong" means harm rather than annoyance. Most teams run one of the three and call it coverage.

Quality is not the only axis either. Latency and cost are quality attributes with users attached to them. A more accurate answer that arrives after the user has given up is a worse product, and any eval that cannot express that will keep recommending the wrong trade.

The test for whether your evals are real is short: name the last release they delayed. If there isn't one, they are decoration, however good the numbers look.

Checklist · · 12 checks

RAG production-readiness checklist

Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.

Checklist · · 24 checks

Working with Claude, practices that hold up

Prompting, agents and tools, Claude Code, evaluation and safety. The habits that make Claude-based systems reliable, as a checklist you can run against your own setup.