The top result from your retriever is the nearest vector, which is often not the best evidence, and the gap between them is where a lot of RAG quality quietly leaks away.

Why first-stage retrieval settles for "near"

Vector search has to scan millions of chunks in milliseconds, so it compares compressed representations, each document embedded once, in advance, with no knowledge of the question it will eventually be asked to answer.

That is enough to gather fifty plausible candidates, but not to pick the three that should shape the answer.

A reranker does the expensive thing instead: it looks at the question and each candidate together and scores how well that specific passage answers that specific question. It is slower by orders of magnitude, far more accurate, and completely impractical across a full corpus, which is exactly why it belongs in second position.

The funnel is the whole design

Cheap and wide first. Expensive and narrow second.

Break the funnel in either direction and you pay for it. Rerank too many candidates and latency and cost balloon for diminishing returns. Skip reranking and raw vector order decides what the model reads, which matters because the model largely trusts whatever you put in front of it. It has no way to know that passage two was a better match than passage one. It will build a confident answer on whatever arrived, in the order it arrived.

Feed a strong model "nearby but wrong" and you get fluent, well-structured, confidently incorrect output. That is a ranking failure wearing a hallucination costume, and teams routinely misdiagnose it as one, then spend a sprint on prompt engineering for a problem no prompt can reach.

Before you add one

Measure the uplift on your own queries. Reranker performance varies sharply by domain. The benchmark number on the model card was not produced on your corpus, your chunk sizes, or your users' phrasing.

Check what reranking cannot fix. If the right document rarely appears in the candidate set at all, reordering is irrelevant. You have a recall problem upstream, and a reranker will simply sort fifty wrong answers very precisely. Diagnose that first: pull twenty failed queries and check whether the correct passage was anywhere in the top fifty. The answer tells you which component to work on.

Closing thought

Retrieval finds candidates and reranking chooses among them. Most systems invest heavily in the finding and leave the choosing to a distance metric that never saw the question.