A prompt is one line in a config file. Someone changes "be concise" to "be concise and friendly", ships it on a Thursday afternoon, and the refusal rate on a safety-sensitive intent quietly halves.

The eval suite would have caught it, but it runs weekly, by hand, when somebody remembers. Only the wiring was missing.

Treat prompts and models like code, because they are

Every input that shapes behaviour deserves a regression test in the release path: the system prompt, the model version and its decoding parameters, the retrieval configuration, the tool definitions. Change any of them and the same suite runs before merge, with no human deciding whether this particular change is worth checking.

That imposes a real constraint. A gating suite has to be fast and cheap enough to survive running on every change, which means keeping it small and deterministic wherever possible. Let the expensive human-reviewed evals run nightly against main instead of blocking a pull request.

An average is an excellent hiding place

Overall score held at 0.86, so it shipped. Underneath, the refusal-required slice fell off a cliff and the sheer volume of easy cases absorbed the difference.

Gate on slices instead of the aggregate. Give each slice its own threshold and let the strictest one hold the release. High-risk slices deserve a rule as blunt as "zero regressions allowed", because averaging is exactly the wrong instinct when the failure is rare and serious.

Classify the failures too, so the trend is legible: fabrication, wrong retrieval, format violation, unsafe compliance, refusal of a legitimate request. A dashboard that reports one number tells an owner nothing about what to do next. A dashboard that reports failure counts by type points at a specific team and a specific fix.

Decide the rollback rule before the canary

Offline evals cannot see everything, so behaviour-changing releases go to a small share of live traffic first. Canaries are only useful if the abort condition was written down beforehand with real numbers on it: which metric, what threshold, over what window, and who is allowed to press the button without convening a meeting.

Pair every such release with a short note recording what changed, which eval covered it, and what the gate said. Six weeks later, when quality has drifted and nobody can reconstruct the order of the prompt edits, that note is the difference between debugging and guessing.

A gate that has never blocked anything has only proven that it is quiet.

Checklist · · 12 checks

RAG production-readiness checklist

Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.

Checklist · · 24 checks

Working with Claude, practices that hold up

Prompting, agents and tools, Claude Code, evaluation and safety. The habits that make Claude-based systems reliable, as a checklist you can run against your own setup.