Chunking strategies
Chunking splits long documents into passages you can retrieve individually. It sounds trivial and is anything but — chunk boundaries decide what can ever be retrieved together, and getting them wrong is one of the most common causes of a RAG system that "just doesn't work."
Why chunk at all?
- Precision. Retrieving a whole 40-page manual to answer one question buries the relevant sentence in noise. Small chunks let you retrieve just the relevant part.
- Embedding limits. Embedding models have a maximum input length; long documents must be split to fit.
- Context budget. You can only fit so many tokens in the LLM prompt. Smaller chunks mean you can include more different passages.
The core trade-off: size
Chunk size is a dial between two failure modes:
- Too small → each chunk lacks context; meaning gets fragmented; a single idea is split across chunks and none retrieves well.
- Too large → chunks contain multiple topics; the embedding becomes a blurry average; irrelevant text rides along and dilutes precision.
There is no universal best size. Common starting points are 200–500 tokens with 10–20% overlap, then tune against a real evaluation set. Overlap — repeating a bit of text between adjacent chunks — prevents a sentence that straddles a boundary from being lost.
The main strategies
1. Fixed-size
Split every N tokens (or characters) with some overlap. Dead simple and fast, but blind to meaning — it will happily cut a sentence or table in half. A fine baseline, rarely the best choice.
2. Recursive / structural
Split on a priority list of separators — paragraphs, then sentences, then words — falling back only when a piece is still too big. This respects natural boundaries far better than fixed-size and is the sensible default for prose. For Markdown or code, split on structural markers (headings, functions) so chunks align with document structure.
3. Semantic
Embed sentences and start a new chunk when the topic shifts (a drop in similarity between consecutive sentences). Chunks follow meaning rather than length, which improves coherence — at the cost of extra embedding calls during indexing.
4. Sentence-window
Embed and retrieve on single sentences for pinpoint matching, but when a sentence is retrieved, expand it to include a window of surrounding sentences before handing it to the LLM. You get precise retrieval and rich context. This is the "small-to-big" idea.
5. Parent-document (small-to-big)
Index small child chunks for precise matching, but return their larger parent chunk (or whole section) to the LLM. Match on specifics, generate on context. Widely used and effective.
6. Agentic / LLM-assisted
Use an LLM to decide boundaries or to write a short contextual summary for each chunk. Highest quality, highest cost — reserve it for high-value corpora. (See Contextual retrieval for a strong variant.)
| Strategy | Precision | Cost | Good for |
|---|---|---|---|
| Fixed-size | Low | Very low | Baselines, uniform text |
| Recursive / structural | Medium | Low | Most prose — the default |
| Semantic | High | Medium | Mixed-topic documents |
| Sentence-window | High | Medium | Fact lookup needing context |
| Parent-document | High | Medium | General-purpose, robust |
| Agentic / contextual | Highest | High | High-value corpora |
Recursive splitting, sketched
function recursiveSplit(text, maxLen, overlap) {
const separators = ["\n\n", "\n", ". ", " "]; // coarse -> fine
function split(chunk, depth) {
if (chunk.length <= maxLen) return [chunk];
const sep = separators[depth] ?? "";
const parts = sep ? chunk.split(sep) : [chunk];
const out = [];
let buf = "";
for (const p of parts) {
const candidate = buf ? buf + sep + p : p;
if (candidate.length > maxLen && buf) {
out.push(buf);
buf = p;
} else {
buf = candidate;
}
}
if (buf) out.push(buf);
// Any piece still too big: recurse with a finer separator.
return out.flatMap(c => (c.length > maxLen ? split(c, depth + 1) : [c]));
}
return withOverlap(split(text, 0), overlap);
}