The generic judge prompt is one of the most quietly expensive artefacts in AI engineering. "Rate this response from 1 to 5 for quality." It returns a 4 for nearly everything, correlates suspiciously well with length and confidence, and nobody in the room can say what a 4 actually means.
Meanwhile the real failure in that response was a fabricated policy figure. The judge gave it a 4, because it was beautifully written.
Do not ask a model what a function can answer
A large share of what teams route to an LLM judge is deterministic. Is the JSON valid against the schema? Are the required fields present? Is every cited document ID actually in the retrieved set? Does the output contain a phone number in a context where it must not?
Write those as code. They run in milliseconds, cost nothing, stay stable across model versions, and when they fail they tell you precisely what broke. Reserve model judgment for the things that genuinely require reading comprehension.
Split the rubric, then calibrate it
Factuality, usefulness, safety and format are four different questions with four different remediations and often four different owners. Collapsed into a single score they cancel each other out, and a charming fabrication lands in the same band as a slightly blunt answer that happens to be correct.
Score them separately, each against criteria specific enough that two reasonable people would agree. "Every claim is supported by a cited passage" is scoreable. "High quality" is not.
Pairwise comparison is worth reaching for when absolute scores keep clustering in the middle. Asking which of two answers is better is a far easier question than assigning a number, for humans and models alike, and ranking a candidate against current production behaviour is usually the decision you actually need.
Then calibrate. Take a few hundred examples your reviewers have already judged, run the judge across them, and measure agreement. A judge that matches your humans seventy per cent of the time is measuring itself more than your product. Re-run that check whenever the judge model, the judge prompt or the rubric changes, because all three drift and only one of them announces it.
Route disagreements to a person rather than averaging them away. Those cases are where the rubric is underspecified, and every one you resolve improves every run that follows.
Closing thought
The test of a rubric is whether the number it produces, on the day that number drops, tells someone specific which thing to go and fix.
More on these topics
Checklist · · 12 checks
RAG production-readiness checklist
Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.
Checklist · · 24 checks
Working with Claude, practices that hold up
Prompting, agents and tools, Claude Code, evaluation and safety. The habits that make Claude-based systems reliable, as a checklist you can run against your own setup.
Explainer · · 2 min read
Real users ask questions your test set never imagined
Offline evals cover the questions you thought of. Production tells you the ones you did not.
Discussion