A support assistant answers a billing question by quoting the refund policy. The quote is accurate, the answer is fluent and correctly cited. It is also wrong, because an exception to that policy sits in the next paragraph of the same document and the assistant did not mention it.
The team fixes the prompt, then adds a reranker, then tries a better model. None of it works, because the exception never reached the model. It was cut into a different chunk during ingestion eight months ago by a script nobody has opened since.
That pattern runs through the whole of retrieval. Retrieval is the only stage of a RAG system that can add information. Reranking reorders, prompting frames, the model reasons. All three operate on a candidate set that retrieval already decided, and none of them can put back something that was never returned.
Which means: retrieval sets the ceiling on everything downstream, and that ceiling is fixed long before any user types a question.
Here is where that ceiling actually gets set.
The document gets cut before anyone asks anything
Chunking is the least interesting topic in RAG and one of the highest-leverage. It happens once, early, usually as a default copied from a tutorial, and it silently caps everything built afterwards.
The damage from a bad boundary is not that chunks look untidy. It is that a claim gets separated from the thing that qualifies it. A rule loses its exception. A figure loses its unit or its date. A table loses the header row that explains the columns, leaving a grid of numbers about nothing. What comes back is individually plausible and collectively misleading, and nothing flags it, because an incomplete chunk does not know it is incomplete.
The bar worth holding to is simple: a chunk should make sense to a person reading it with no surrounding document. If a human cannot follow it in isolation, the model has no advantage over that human.
The half most teams skip is that changing your chunking strategy changes what is findable across the entire corpus. Answers that worked break, answers that were broken start working. That is release-sized blast radius, and it deserves a fixed evaluation set run before and after, with someone actually reading the diff.
Semantic search gets you to the neighbourhood, not the house
Embeddings turn text into coordinates, positioned so distance approximates similarity of meaning. That is what lets a user ask about "money back" and reach a document that only ever says "refund".
The production problem is a quiet substitution. Embeddings measure similarity. Teams use them as though they measured relevance. Those two agree often enough to be convincing and diverge often enough to hurt. "What is our refund policy" and "customers furious about refunds" sit close together in embedding space. One is a policy lookup, one is a complaints analysis, and the score will not tell you which you got.
The sharpest version is negation. "Eligible" and "not eligible" embed remarkably close together, because a negating word barely moves the coordinates. Any domain where a "not" flips the outcome, which is to say eligibility, compliance, safety and permissions, needs a signal that is not purely semantic.
Two traps follow. General-purpose embedding models are trained on general-purpose text, so your internal product names and acronyms may be positioned arbitrarily. And a similarity score of 0.87 looks like a probability of correctness. It is a distance in a space you did not design, and a threshold that looked sensible on sample data can mean nothing on real queries.
The index stops matching what your organisation knows
A vector store gets marketed as memory for your AI. Memory sounds like something that maintains itself. A database very much does not.
The failure here is decay rather than drama. When a document is retracted, corrected or superseded, do its vectors actually leave the index? In a surprising number of systems, no. The old vectors stay perfectly retrievable, still answering questions with content the company formally withdrew. Under a deletion request, that stops being an engineering problem and becomes a legal one.
The quieter versions matter too. Metadata filters interact badly with approximate nearest-neighbour search, so a query scoped to one department can return worse results than the unfiltered version, because the filter shrinks the candidate pool the index was tuned for. Without scheduled re-indexing, content drifts from reality and nothing alerts you, because stale vectors return results just as confidently as fresh ones.
The unfashionable observation attached to all this: at small scale you may not need a dedicated vector database. Postgres with pgvector covers a wide middle ground in a system your team already backs up. The dedicated store earns its place when scale, latency or filtering genuinely demand it.
Picking one search method is picking which users to fail
Search for an error code semantically and you may get a thoughtful paragraph about error handling in general. Search for "why do customers stop paying" by keyword and you get nothing, because no document contains that sentence.
Keyword matching is fast, precise, explainable, and completely indifferent to the fact that "cancel" and "terminate" mean the same thing to your users. Semantic search handles paraphrase beautifully and is unreliable on exactly the strings where one character changes everything: part numbers, invoice references, policy names.
The move that settles this is empirical and takes about an hour. Pull a day of real production queries and sort them into the ones that hinge on an exact string and the ones that describe an intent. Almost every real query stream has serious volume in both piles, and the ratio is rarely what the team assumed in the design meeting.
Then notice why this persists. When keyword-only search misses a paraphrase, the user sees no results and rephrases or leaves. When semantic-only search mangles an invoice number, the user sees a plausible wrong document and may never realise. Neither looks like an architecture problem. Both look like one unlucky answer. The gap is visible only in aggregate, which is precisely why you have to look in aggregate.
Running both is easy, merging both is the actual work
Hybrid search declines the choice: run both retrievers, combine the results, cover each blind spot with the other's strength. It is the boring answer and usually the right one.
What "combine" hides is a real decision. Keyword scores and vector similarities live on completely different scales, so something has to reconcile them, typically Reciprocal Rank Fusion using positions instead of raw scores. That fusion step now decides what the model sees, and its defaults were tuned on someone else's corpus. It is a first-class part of retrieval quality and needs testing like one.
Three things worth settling before shipping. Decide deliberately what happens when the same chunk arrives from both paths, since two independent signals agreeing is real evidence. Budget latency across two searches plus fusion plus a likely reranker, and measure p95 rather than average. And log which path won, because a merged result with no per-path attribution is undebuggable.
Then make it earn its complexity on your data. Sometimes the gain is dramatic, and sometimes your corpus is almost entirely natural language and semantic-only is fine. That is a measurement, not a preference.
Relevance and sensitivity rise together
The last one turns a quality problem into an incident.
Your system finds the perfect document for the question. It is the executive compensation file, and the person asking is an intern. Retrieval was accurate, ranking was sensible, the answer was well grounded, and you have a serious problem.
Semantic search has an uncomfortable property: sensitive documents surface exactly when they are most relevant to whoever is asking. The moment someone asks about salaries is the moment the compensation file ranks first. Relevance and sensitivity correlate, so this happens routinely rather than at the edges. The system is working as designed, aimed at the wrong audience.
There are two places the rule can live and they are not equivalent. In the prompt, "do not reveal salary information" is a request to a probabilistic system that has already read the salary information. In the query, the retrieval call is scoped by the requesting user's identity, groups and clearance, and the file never enters the candidate set. Only the second is a control. If your security story is that the model was instructed not to, you do not have a security story.
Two rules follow. Mirror your existing permissions rather than rebuilding them, because a parallel copy of your access model will drift from the real one and drift silently. And fail closed, because degrading to unfiltered retrieval when the permission service times out is an availability decision that quietly becomes a disclosure decision.
Six failures, none of which throws an error
Six posts, one argument: retrieval decides what your system is capable of knowing, and every decision that sets it is made early, is expensive to reverse, and fails silently.
Look at what these six share. None of them throws an error. A bad chunk boundary, an overtrusted similarity score, a stale index, a missing search signal, an untested fusion step, an unscoped query: not one produces an exception or an alert. Each produces a confident answer built on the wrong evidence, which at a glance is indistinguishable from a confident answer built on the right evidence.
Three of them behave like schema migrations rather than configuration. Change your chunking, change your embedding model, or let your index drift, and you have altered what is findable across the whole corpus. Teams that treat these as parameters hear about the regression from a user.
So retrieval quality is not one dimension. Evidence can be relevant and incomplete, relevant and stale, relevant and unauthorised. Measuring retrieval separately from generation, against a fixed set of questions you rerun on every index change, is what turns this from a series of invisible failures into something you can actually see.
The six posts in full
Each of these is a short standalone post on my blog, with the specifics, the failure modes and what to actually do:
- Why chunk boundaries shape answer quality
- Why embeddings are useful and easy to overtrust
- Why vector stores are infrastructure, not magic memory
- Why keyword search and semantic search both matter
- Why hybrid search is often the practical default
- Why relevant data can still be unauthorized data
They are the third chapter of 60 Days of Production AI Systems, a series on building AI that survives real users. If you would rather start from the beginning, Start Here lays out all sixty in ten chapters.
More on these topics
Checklist · · 12 checks
RAG production-readiness checklist
Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.
Failure story · · 1 min read
The retrieval cache that served stale policies
A well-meaning cache in front of retrieval kept answering from last quarter's HR policy for eleven days.
Deep dive · · 6 min read
Production RAG does not fail loudly, and that is the whole problem
Six days of notes on operating retrieval systems after launch, where nearly every real failure arrives dressed as a good answer.
Discussion