RAG Handbook retrieval-augmented generation Bipin Singh
Advanced RAG

Contextual retrieval

2 min readChapter 16 of 26By Bipin Singh

When you split a document into chunks, each chunk loses the context of the whole. A chunk saying "the limit is 500 requests per minute" doesn't say which product or plan it refers to — so it retrieves poorly for a specific question. Contextual retrieval fixes this by prepending situating context to each chunk before embedding it.

The lost-context problem

Consider a chunk: "Revenue grew 3% over the previous quarter." Which company? Which quarter? A user asking "how did Acme perform in Q2 2024?" may never retrieve it, because the chunk itself contains none of those anchors. Splitting created an orphan.

The fix

Before embedding, use an LLM to generate a short, chunk-specific blurb that locates the chunk within its document, and prepend it. The example above becomes: "This chunk is from Acme's Q2 2024 report; it discusses quarterly revenue. Revenue grew 3% over the previous quarter." Now it embeds — and retrieves — with the right anchors.

// Index time: contextualize each chunk against its parent document.
for (const chunk of chunks) {
  const situating = await llm.complete(
    "Give a short sentence situating this chunk within the document, " +
    "to improve search retrieval. Answer with the sentence only.\n\n" +
    "Document title: " + doc.title + "\n" +
    "Chunk: " + chunk.text
  );
  chunk.embedText = situating + "\n\n" + chunk.text; // embed this
  chunk.text = chunk.text; // still store the original for the prompt
}

Apply it to both retrieval channels

Prepend the context before generating the embedding and before building the keyword (BM25) index. Improving both dense and sparse retrieval compounds with hybrid search and reranking for a substantial reduction in failed retrievals.

Managing the cost with prompt caching

Contextualising every chunk means an LLM call per chunk at index time — potentially expensive on a large corpus. Prompt caching makes this affordable: put the full document in the cached portion of the prompt so you pay to process it once, then reuse it cheaply across all its chunks. Index-time cost is a one-off; the retrieval gains are permanent.

Key idea
Contextual retrieval, hybrid search, and reranking stack. Each attacks a different failure mode — missing context, missing exact terms, and coarse ranking — so together they compound rather than overlap.
Bipin Singh
Written by Bipin Singh

Senior Full-Stack Engineer · AI & AWS. I build production RAG and LLM systems for enterprises.

Work with me