Somebody wrote a loader in a notebook, ran it once, and the index has been sitting there ever since. Ask which documents failed to parse and you get a shrug. Ask how many of the 12,000 indexed files still exist in the source system and you get a longer one.
People call that a pipeline. It is a snapshot with a good launch narrative.
Parsing decides your ceiling
Whatever structure you destroy at parse time cannot be recovered downstream by a better embedding model or a cleverer reranker.
PDFs are the usual offender. A financial table flattened into one line of drifting numbers is worse than useless, because it retrieves confidently and answers wrong. Headings collapsed into body text remove the section boundaries chunking depends on. Page references vanish, and with them any hope of a citation a user can verify.
So normalisation gets its own stage and its own tests. Convert to a structured intermediate form, keep the heading hierarchy, keep table cells as cells, keep page and section anchors attached to the text. Chunking then reads structure instead of guessing at it from whitespace.
Metadata extraction belongs here too, while the document still has context around it. Title, author, effective date, document type, section path. Pulled later from a bare chunk, most of it is unrecoverable.
Failures need a destination
Every ingestion run produces documents that fail. The only question is whether anyone finds out.
A production pipeline separates transient failures, which retry with backoff, from poison documents, which go to a quarantine queue with the error attached and stop being retried forever. Both get counted. If three percent of a corpus silently fails to ingest, the assistant will confidently report that no policy exists on a topic documented inside that three percent.
Treat the run like any ETL job: documents in, documents indexed, documents quarantined, duration, and an alert when the delta moves.
Updates and deletes are the hard half
First ingestion is the easy case, and it is the only case most teams test.
Re-ingestion needs idempotency, usually a stable document ID plus a content hash, so an unchanged file does not churn the index and a changed one replaces its old chunks instead of joining them. Without that you accumulate two versions of the same policy, and retrieval picks whichever one happens to embed better.
Deletes need to propagate on a clock you can state out loud. That means tombstones, a reconciliation sweep comparing indexed IDs against the source of truth, and a metric for how long removal actually takes end to end.
Closing thought
RAG quality is mostly decided upstream of the model, in a pipeline that either gets engineered like a data product or gets rediscovered during an incident.
More on these topics
Checklist · · 12 checks
RAG production-readiness checklist
Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.
Failure story · · 1 min read
The retrieval cache that served stale policies
A well-meaning cache in front of retrieval kept answering from last quarter's HR policy for eleven days.
Deep dive · · 6 min read
Production RAG does not fail loudly, and that is the whole problem
Six days of notes on operating retrieval systems after launch, where nearly every real failure arrives dressed as a good answer.
Discussion