RAG Handbook retrieval-augmented generation Bipin Singh
The core pipeline

Chunking strategies

3 min readChapter 05 of 26By Bipin Singh

Chunking splits long documents into passages you can retrieve individually. It sounds trivial and is anything but — chunk boundaries decide what can ever be retrieved together, and getting them wrong is one of the most common causes of a RAG system that "just doesn't work."

Why chunk at all?

The core trade-off: size

Chunk size is a dial between two failure modes:

There is no universal best size. Common starting points are 200–500 tokens with 10–20% overlap, then tune against a real evaluation set. Overlap — repeating a bit of text between adjacent chunks — prevents a sentence that straddles a boundary from being lost.

Key idea
Chunk to a single coherent idea. The goal is a passage that stands on its own: specific enough to match a real question, complete enough to answer it.

The main strategies

1. Fixed-size

Split every N tokens (or characters) with some overlap. Dead simple and fast, but blind to meaning — it will happily cut a sentence or table in half. A fine baseline, rarely the best choice.

2. Recursive / structural

Split on a priority list of separators — paragraphs, then sentences, then words — falling back only when a piece is still too big. This respects natural boundaries far better than fixed-size and is the sensible default for prose. For Markdown or code, split on structural markers (headings, functions) so chunks align with document structure.

3. Semantic

Embed sentences and start a new chunk when the topic shifts (a drop in similarity between consecutive sentences). Chunks follow meaning rather than length, which improves coherence — at the cost of extra embedding calls during indexing.

4. Sentence-window

Embed and retrieve on single sentences for pinpoint matching, but when a sentence is retrieved, expand it to include a window of surrounding sentences before handing it to the LLM. You get precise retrieval and rich context. This is the "small-to-big" idea.

5. Parent-document (small-to-big)

Index small child chunks for precise matching, but return their larger parent chunk (or whole section) to the LLM. Match on specifics, generate on context. Widely used and effective.

6. Agentic / LLM-assisted

Use an LLM to decide boundaries or to write a short contextual summary for each chunk. Highest quality, highest cost — reserve it for high-value corpora. (See Contextual retrieval for a strong variant.)

StrategyPrecisionCostGood for
Fixed-sizeLowVery lowBaselines, uniform text
Recursive / structuralMediumLowMost prose — the default
SemanticHighMediumMixed-topic documents
Sentence-windowHighMediumFact lookup needing context
Parent-documentHighMediumGeneral-purpose, robust
Agentic / contextualHighestHighHigh-value corpora

Recursive splitting, sketched

function recursiveSplit(text, maxLen, overlap) {
  const separators = ["\n\n", "\n", ". ", " "]; // coarse -> fine
  function split(chunk, depth) {
    if (chunk.length <= maxLen) return [chunk];
    const sep = separators[depth] ?? "";
    const parts = sep ? chunk.split(sep) : [chunk];
    const out = [];
    let buf = "";
    for (const p of parts) {
      const candidate = buf ? buf + sep + p : p;
      if (candidate.length > maxLen && buf) {
        out.push(buf);
        buf = p;
      } else {
        buf = candidate;
      }
    }
    if (buf) out.push(buf);
    // Any piece still too big: recurse with a finer separator.
    return out.flatMap(c => (c.length > maxLen ? split(c, depth + 1) : [c]));
  }
  return withOverlap(split(text, 0), overlap);
}
Tip
Don't pick a chunking strategy in the abstract — pick a couple, run them against a set of real questions with known answers, and measure retrieval recall. Chunking is the single highest-leverage knob you can tune empirically.
Bipin Singh
Written by Bipin Singh

Senior Full-Stack Engineer · AI & AWS. I build production RAG and LLM systems for enterprises.

Work with me