What we built
A query-level cache in front of retrieval to cut cost and latency. Cache key: a hash of the normalised question. Hit rate was excellent.
What broke
The policy corpus was re-indexed after a change to parental leave rules. The cache did not know. For eleven days, anyone who asked a question someone else had already asked got the old answer, with a confident citation to a document that no longer said that.
Why nobody noticed
Retrieval recall on the golden set stayed perfect, because the golden set was built from the old corpus. Latency dashboards looked better than ever. There was no metric for "answer freshness".
What we changed
Every index build writes a version. The version is part of the cache key. A re-index invalidates everything, which is the correct trade.
We also added a freshness metric: the age of the newest source passage behind each answer, plotted against corpus age.
The lesson
A cache in front of a knowledge base stores answers that were true for one corpus version. Put that version in the key.
Key takeaways
- Cache keys must include the corpus version as well as the query.
- Staleness is invisible without a freshness metric on answers.
- The fix was ten lines, but detecting the problem took eleven days.
More on these topics
Checklist · · 12 checks
RAG production-readiness checklist
Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.
Article · · 1 min read
What an AI gateway actually costs to run
The operational bill for one gateway in front of a dozen tool servers, including the costs nobody budgets for.
Explainer · · 2 min read
Nothing broke, and the system is still getting worse
Drift is quiet by construction. SLOs and a standing review are what make it audible.
Discussion