Chunking: The Irreversible Decision
ONELINE
Chunking is the one retrieval decision no later model can rescue, because information destroyed at index time is simply not there at query time.
Documents are too long to embed whole. So you cut them up. Each piece gets embedded, indexed, and retrieved independently.
This step gets treated as plumbing — pick 512 tokens, move on. It is not plumbing. It is the decision that determines the ceiling on everything downstream, and unlike almost every other choice in this book, you cannot fix it later without rebuilding the index.
A better embedding model cannot recover a sentence you split in half. A better reranker cannot re-join a table separated from its caption. A better LLM cannot infer the pronoun reference you cut away.
The tension
Every chunking decision trades two things against each other:
| Small chunks (~100 tokens) | Large chunks (~1,000 tokens) | |
|---|---|---|
| Precision | High — tightly on topic | Low — meaning diluted |
| Context | Poor — broken fragments | Rich — surrounding detail |
| Storage | High — many vectors | Low — few vectors |
| Search latency | Lower per vector, more vectors | Higher per vector, fewer |
The rule underneath it:
Smaller is better for finding. Larger is better for thinking.
A 100-token chunk is one clear idea, so its embedding is sharp and it matches queries precisely. But hand it to the model and it may be a fragment — a sentence whose subject was in the previous chunk.
A 1,000-token chunk gives the model plenty to reason with. But its embedding is an average of a dozen ideas, so it matches everything vaguely and nothing well. This is semantic dilution, and it is the main reason naive large-chunk systems retrieve mush.
The strategies, in order of sophistication
Fixed-size splitting
Cut every N tokens. Fast, predictable, and it will slice through the middle of sentences, tables, and code blocks.
Use it as a baseline to measure against, not as a destination.
Overlap is the cheap patch: let consecutive chunks share 10–20% of their text, so a sentence cut at a boundary survives intact in its neighbour. It raises storage and adds duplicate retrievals, and it is worth it.
Recursive structure splitting
Split on the document’s own structure, in priority order: paragraphs first, then sentences, then words — descending only when a piece is still too big.
This respects what the author already told you about how the content groups. It is the sensible default for prose, and the point where most teams should start.
Semantic chunking
Embed each sentence, then start a new chunk wherever consecutive sentences become dissimilar — cutting at topic boundaries rather than structural ones.
Better boundaries, at the cost of embedding every sentence at index time. Worthwhile for corpora where structure is unreliable — transcripts, scanned documents, anything without clean headings.
Parent-child (hierarchical)
This is the one that resolves the tension rather than trading within it, and it is the production standard.
Split the document into large parent chunks of around 1,500 tokens. Split each parent into small child chunks of around 300. Then:
- Index only the children. Small, precise, easy to match.
- When a child matches a query, return its parent to the model.
You search with the small chunk’s precision and generate with the large chunk’s context. The tension does not go away — you just stop having to pick a side.
ASIDE
This pattern recurs: search cheap and narrow, deliver rich and wide. Chapter 9 applies the same shape to ranking.
Content dictates the strategy
The right boundary depends entirely on what you are cutting.
| Content | Split on | Never split |
|---|---|---|
| Prose, documentation | Headings, then paragraphs | Mid-sentence |
| Code | Function and class boundaries | Inside a function body |
| Tables | Keep whole; repeat the header per chunk | Rows from their headers |
| Transcripts | Speaker turns, topic shifts | Mid-utterance |
| Legal, policy | Clause and section numbers | Clause from its conditions |
| Q&A, FAQ | One question and answer per chunk | Question from answer |
WATCHOUT
Tables are where naive chunking does its worst damage. Split a table from its header and every row becomes uninterpretable — the model sees 4.2 | 17 | yes with no idea what those columns are.
The same applies to any content where meaning lives in structure rather than prose: split a code function from its signature, or a clause from its conditions, and you have indexed something that is worse than nothing, because it will still match queries and then mislead.
Contextual chunking
There is one more move, and it is the highest-leverage recent idea in this area.
The problem: a chunk reading “The policy does not apply in this case” is useless in isolation. Which policy? Which case?
The fix: before embedding, prepend a short generated description situating the chunk in its document.
original chunk:
"The policy does not apply in this case."
contextualised chunk:
"From the 2025 Refund Policy, section 4, digital goods:
The policy does not apply in this case."
Now the chunk is self-describing, and it matches queries about refund policies for digital goods — which the original never could.
The cost is one LLM call per chunk at index time. With prompt caching over the parent document, this is far cheaper than it sounds, and it reliably produces one of the largest single improvements available to a struggling RAG system.
Choosing, and then proving
A reasonable default: recursive splitting into ~300-token children with ~15% overlap, grouped under ~1,500-token parents, with contextual prefixes if quality matters more than index cost.
But the real advice is to stop treating this as a thing to be guessed. Build a set of question-and-expected-source pairs, then measure what fraction of the time the right source is retrieved under each strategy. Chunking is cheap to A/B at index time and ruinously expensive to get wrong at query time.
The other thing you cannot change incrementally
Chunking is not the only index-time decision with no partial fix. The embedding model is the other one, and it fails in a nastier way.
Embeddings from two different models are not comparable. Not “slightly worse together” — meaningless together. Two models trained separately produce unrelated coordinate systems, like comparing a grid reference to a postcode.
WATCHOUT
Changing or upgrading your embedding model invalidates your entire index. Every document must be re-encoded before any query uses the new model. There is no partial migration: a corpus half in the old space and half in the new returns confident nonsense, because the arithmetic still works — it just means nothing.
Plan for it. Version the model id alongside the vectors, build the new index in parallel, and cut over atomically. Teams that skip this discover it as a slow, unexplainable relevance collapse.
The same applies more subtly to model versions, and to changing the normalisation. If the vectors were produced differently, they do not belong in the same index.
The architectural lesson is worth stating as a rule, because it is not obvious until it bites: an index is not a store of your documents. It is a store of one model’s opinion about your documents. Treat it that way and the design follows. Stamp the index with the model id, its version, the dimension count and whether vectors were normalised; refuse a write whose stamp disagrees; and make “rebuild the index” a routine, scheduled, measured operation rather than an emergency. A pipeline that can rebuild from source on demand turns a migration into an afternoon. One that cannot has quietly made its embedding model permanent.
The difference from chunking is worth stating precisely, because it is the reason this chapter is the one named irreversible. Re-embedding is expensive but it recovers everything — the source documents are untouched, so the new vectors are as good as the new model. Re-chunking is expensive and recovers only what your old boundaries did not already destroy.
What to carry forward
Chunking sets the ceiling. Small for finding, large for thinking, parent-child to get both. Respect the content’s own structure, never separate a table from its header, and consider giving each chunk enough context to describe itself. And version the embedding model beside the vectors, because that migration has no incremental path either.
Next: what to do with the candidates once you have them.
RECALL
- Why is chunking harder to fix later than the choice of embedding model?
- State the small-versus-large trade-off in one sentence each way.
- Explain parent-child chunking. What is indexed, and what is returned?
- What is semantic dilution, and which chunk size causes it?
- Your corpus is financial tables. Name two chunking rules you would enforce.
- A chunk reads “This does not apply to enterprise customers.” Why will it retrieve badly, and what would you do at index time?
- A colleague proposes upgrading the embedding model “gradually, as documents are updated.” Explain precisely what will go wrong.