A retrieval bug in development is easy. The system returns nothing, or it returns something so obviously wrong that you catch it before lunch.

In production the same class of bug returns a fluent paragraph that cites a real document and happens to be wrong. A support assistant quotes a discount that ended two quarters ago. Retrieval succeeded. The citation resolves. The groundedness score is excellent. Every panel on the dashboard is green. You find out from a customer, or from legal, or from the agent who noticed the bot repeating something that stopped being true in March.

That is the shape of nearly every failure in this chapter. Once a RAG system is live, its failure modes stop announcing themselves, and the operational work is almost entirely about converting silent failures into visible ones. Here is where the silence comes from.

The question that looks like a document question and is not

"What is our refund policy?" and "where is my refund?" arrive in the same chat box, phrased almost identically, and are answered by completely different infrastructure.

The first lives in prose. The second lives in a database and changes faster than any index can track. Inventory, order status, ticket counts, balances, entitlements: none of these belong to a document pipeline, and a document pipeline asked for them will not refuse. It finds the closest-matching text, usually a status report from a quarter ago, and answers with total composure.

So routing has to be explicit and logged. "It queried the wrong system" and "it queried the right system badly" are different bugs with different fixes, and you cannot tell them apart afterwards unless the decision was recorded at the time.

Safety comes next. A model writing free-form SQL against production is an incident generator with a natural-language interface. Five narrow typed tools beat one general query tool every time: easier to validate, easier to log, much harder to misuse.

Then decide what happens when the live system is down. "The database timed out" must never quietly become "answered from a cached document." That substitution is invisible by construction, which makes it exactly the kind of thing worth alerting on.

What the pipeline discarded before anyone asked a question

Audit what your organisation actually knows and very little of it is clean prose. It is tables buried in PDFs, numbers on dashboards, diagrams in decks, screenshots stapled to old tickets.

The failure is subtle, because most pipelines do process these formats. They just process them badly and produce plausible text either way. A table flattened into a paragraph loses the relationship between its rows and columns. A chart summarised as "a bar chart showing quarterly performance" has thrown away every number a user might ask about. Nothing throws an error. The degraded text gets embedded, indexed and retrieved with full confidence, and the answer is built on a version of your data that has had its structure removed.

Two things follow. Prioritise by value rather than by novelty: image search is the exciting part of multimodal retrieval, but tables are the profitable part, because tables are where the numbers people ask about actually live. And keep a pointer back to the original artifact. A model-written description of a chart is a retrieval aid, not evidence. When a reviewer asks where a number came from, they should see the chart, not another model's paraphrase of it.

Correctness has an expiry date

Staleness is the most under-managed failure mode in production RAG, and the reason is structural. Every other failure has a symptom. A retrieval miss produces a visibly thin answer. A permissions leak triggers an incident. An outage pages somebody. Stale content produces a confident, well-sourced, correct-looking answer to a question whose truth has moved, and nothing in the system knows anything is wrong.

The fix is a posture change. An archive preserves; a newsroom updates and retires. RAG needs the second one. Ingestion needs an owner and monitoring, and the alert that matters is documents-ingested dropping to zero, because a job that succeeds while processing nothing is the quieter version of the same outage. Superseded content has to be expired rather than left in the index to compete, since version two may well be the better semantic match and retrieval will sometimes prefer it. And time-sensitive answers should show their vintage: "as of the March 2026 policy" is the cheapest trust mechanism in RAG, routinely cut for feeling like clutter, and it is the difference between a claim a user can check and one they have to take on faith.

One number is worth measuring that almost nobody tracks: how long does it take for a changed document to change your system's answers? Index lag plus cache plus embedding queue adds up, and the total is usually longer than anyone assumed. If that window is a day, say so in the product. If it is a month, it is a roadmap item.

The gap where nothing currently happens

Ask what your system does when retrieval comes back weak and the honest answer, for most systems, is that it does exactly what it does when retrieval comes back strong. Between retrieval and generation there is no gate. Whatever came back flows into the model, and the model does what models do, which is write something fluent with it.

Corrective retrieval puts a checkpoint in that gap. Before generating, grade the evidence: relevant, sufficient, consistent, current enough. If it passes, proceed. If not, rewrite the query and search again, widen to another source, or say plainly that nothing solid was found. The grader is judging relevance rather than writing prose, so a small fast model does it well. Cap the retries at two, because an uncapped correction loop is a cost incident that looks reasonable at every individual step, and log every correction, because the queries that needed a second pass are a precise map of where first-pass retrieval is weak.

The part teams resist is refusal, which feels like admitting the product does not work. Users read it differently. They forgive "I could not find a reliable answer" almost entirely, and they do not forgive being confidently misled. Better still, refusal rate by topic is a free roadmap: a cluster of refusals around one product area is the system telling you which documentation to write next. A pipeline that always answers cannot report its own gaps.

Self-checking only counts when it is anchored and consequential

"Have the model check its own work" is either one of the highest-value additions to a RAG pipeline or a pure waste of tokens, and the difference is what the check is anchored to.

Ask a model whether an answer is good and you get a thoughtful paragraph concluding that it is, on balance, quite good. That is the model grading its own prose with the same intuitions that produced it. Ask instead whether source two supports claim three and the operation has changed entirely. It is verification against an external artifact, and it works because verification is genuinely easier than generation. The same model that invented a detail can often catch the invention when made to line the sentence up against the retrieved text.

Then wire in consequences, because the usual way this pattern burns money is that the critique runs, reads well, and nothing downstream consumes it. An unsupported claim should be stripped or marked. Several should trigger another retrieval pass. An unsupported core claim should block the answer. If you cannot name what changes when the verdict comes back negative, do not add the step yet.

One quality score tells you nothing you can act on

"The RAG system is bad" is a diagnosis-free diagnostic. It says something is wrong and nothing about where, so teams fix the layer they can see, which is the prompt. Weeks of prompt iteration later, retrieval was broken the whole time.

A RAG pipeline is several systems wearing one interface, and they fail independently. Did the knowledge enter the index. Did the right passage make the candidate set. Did it rank high enough to be read. Did the answer stick to it. Can a reviewer verify each claim.

Diagnose upstream to downstream, because upstream failures disguise themselves as downstream ones. If the correct passage was never retrieved, the symptom looks exactly like hallucination: the model had nothing right to work from and produced something plausible instead. A large share of "the model is making things up" tickets resolve to a recall number nobody was measuring.

The setup is small. Fifty labelled questions, each with the passage that should answer it and a note on what a good answer contains. An afternoon of work, run on every meaningful change, and re-chunking stops being a gamble.

Wrong answers that look exactly like right ones

Six ideas, one argument: a production RAG system's real risk is the wrong answers that look exactly like right ones, and every control worth building exists to make that difference observable.

Routing that is logged. Evidence that keeps a pointer to its original. Timestamps on claims that expire. A gate that can decline. A verification step with authority over what ships. Metrics that name the failing layer rather than the failing system.

None of these makes the model smarter. Every one of them makes the system legible enough to operate.

Read the detail

Each of these is a short standalone post on my blog, with the specifics, the failure modes and what to actually do:

They are the fifth chapter of 60 Days of Production AI Systems, a series on building AI that survives real users. If you would rather start from the beginning, Start Here lays out all sixty in ten chapters.

Checklist · · 12 checks

RAG production-readiness checklist

Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.