- Day 13: Why chunk boundaries shape answer quality
- Day 14: Why embeddings are useful and easy to overtrust
- Day 15: Why vector stores are infrastructure, not magic memory
- Day 16: Why keyword search and semantic search both matter
- Day 17: Why hybrid search is often the practical default
- Day 18: Why relevant data can still be unauthorized data
Split a document in the wrong place and you can make its answer unfindable.
That is chunking: nobody's favourite topic, and a bigger lever on RAG quality than most model decisions. It also gets made exactly once, early, by whoever wrote the ingestion script, and then silently caps everything built on top of it.
What a bad boundary actually does
A bad boundary does more than make a chunk look untidy. It separates meaning from the thing that qualifies it.
A rule gets split from its exception, so the system confidently states the rule. A number gets split from its unit or its date. A table gets split from the header row that says what the columns are, which turns a grid of figures into a grid of figures about nothing. A "however" ends up in a different chunk from the sentence it was correcting, and the correction simply ceases to exist as far as retrieval is concerned.
Retrieval then returns pieces that are individually plausible and collectively misleading. The model has no way to know something is missing; an incomplete chunk does not announce itself.
Fixed-size is a default, not a decision
Splitting every document into 500-token blocks is fast, uniform, and completely indifferent to what the documents say. It will cut mid-sentence, mid-table, mid-clause, and it will do so at scale.
It is a reasonable starting point but not a strategy, and the tell is that nobody can explain why it is 500 rather than 300 or 800.
Cutting where the meaning cuts (sections, headings, paragraphs, list items, whole table-with-caption units) costs more work at ingestion and pays back on every query afterwards. Keep a policy with its exceptions. Keep a definition with what it defines.
A usable bar: a chunk should make sense to a person reading it alone, with no surrounding document. If it does not, the model has no better chance than the person did.
Re-chunking is a release, not a config change
Changing your chunking strategy changes what is findable across your entire corpus. Answers that worked can break; answers that were broken can start working. That is the same blast radius as swapping the embedding model, and it deserves the same treatment: a fixed set of evaluation questions, run before and after, with the diff actually reviewed.
Teams that treat re-chunking as a parameter tweak find out about the regression from a user.
Closing thought
When a RAG answer quotes the rule and misses the exception sitting directly beneath it, the instinct is to blame the model. Check the boundary first. The exception was probably in the part you never retrieved.
Try it: the chunking lab lets you move the chunk size and overlap and watch an answer split across a boundary.
More on these topics
Checklist · · 12 checks
RAG production-readiness checklist
Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.
Failure story · · 1 min read
The retrieval cache that served stale policies
A well-meaning cache in front of retrieval kept answering from last quarter's HR policy for eleven days.
Deep dive · · 6 min read
Production RAG does not fail loudly, and that is the whole problem
Six days of notes on operating retrieval systems after launch, where nearly every real failure arrives dressed as a good answer.
Discussion