The first eval set is usually forty rows in a spreadsheet, written by one engineer in an afternoon, drawn from queries they happened to remember. It is genuinely useful for about six weeks. Then it starts lying, quietly, in the direction of whatever the team already believes.

What actually belongs in the set

Representative traffic is only the floor. If the set mirrors production volume, it will be ninety per cent easy cases and the average will drown every failure worth finding.

Build it as three deliberate slices and score them separately:

  • Common cases, enough to notice a broad regression
  • Long tail: odd formats, multi-part questions, the languages nobody planned for
  • Adversarial: injection attempts, questions with no answer in the corpus, requests that should be refused

The third slice is the one teams skip, because adding it makes the dashboard look worse. That is precisely its job.

Labels are only as stable as the guide

Two reviewers will label the same borderline answer differently, and neither of them is wrong. They are applying different definitions because nobody wrote one down.

So write the reviewer guide before labelling, with worked examples of a pass, a fail and the ambiguous case that sits between them. Measure agreement between reviewers on a shared subset. Persistent disagreement is a specification bug rather than a people problem, so fix the guide, then relabel.

Version the data and the rules together

The biggest time sink is a comparison across versions. The labelling guidance is tightened in March, the set is relabelled, and in April someone compares the new score to February's as though nothing changed. The model may not have moved at all.

Treat the dataset and its guide as one versioned artefact. Every example carries provenance: where it came from, who labelled it, under which version of the rubric. When a comparison spans a version boundary, say so on the chart instead of hoping nobody notices.

The set should also grow on purpose. Production failures that reached a user get triaged back in as new cases, with the corrected output attached. That way the set reflects your users rather than your assumptions.

Keep it out of the tuning loop

An eval set that has been used for prompt tuning is a training set wearing a costume. Every round where someone read the failures and edited the prompt to fix them has fitted that prompt to that data, and the score has stopped being evidence.

Hold back a slice nobody optimises against and nobody browses. Rotate it rarely and deliberately. It is the only number you can quote outside the team without a caveat.

All of this needs one owner: the guide, the versions, the intake of new failures. Shared custody means nobody notices when it goes stale.

When did your eval set last gain an example that came from a real user complaint?

Checklist · · 12 checks

RAG production-readiness checklist

Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.

Checklist · · 24 checks

Working with Claude, practices that hold up

Prompting, agents and tools, Claude Code, evaluation and safety. The habits that make Claude-based systems reliable, as a checklist you can run against your own setup.