RAG Handbook retrieval-augmented generation Bipin Singh
The core pipeline

Retrieval

2 min readChapter 08 of 26By Bipin Singh

Retrieval is the heart of RAG. Given a query, it decides which passages the model gets to see — and if the right passage isn't retrieved, the answer cannot be right. This page covers how similarity is measured and the search strategies that find the best passages.

Similarity metrics

Search ranks candidates by how "close" their vectors are to the query vector. Three measures dominate:

MetricMeasuresNotes
Cosine similarityAngle between vectorsIgnores magnitude; the most common choice for text
Dot productAngle and magnitudeEquivalent to cosine when vectors are normalised; cheaper
Euclidean (L2)Straight-line distanceSensitive to magnitude; less common for text

Use the metric your embedding model was trained for — models usually specify one. For normalised embeddings, cosine and dot product rank identically.

Top-k

Top-k is how many passages you retrieve. Too few and you risk missing the answer; too many and you pad the prompt with noise (and pay for it). Typical values are 3–10, but the right number depends on chunk size and the model's context budget. Tune it — and consider retrieving many candidates, then reranking down to a few (see Reranking).

Dense vs sparse retrieval

Key idea
Dense and sparse fail in opposite ways. Dense misses exact terms; sparse misses meaning. Combining them — hybrid search — is often the single biggest, cheapest win in real systems.

Hybrid search

Hybrid search runs both a dense and a sparse query and fuses the results. The most robust fusion method is Reciprocal Rank Fusion (RRF), which combines results by rank rather than by raw scores — so you don't have to reconcile incomparable score scales.

// Reciprocal Rank Fusion: merge two ranked lists by position.
function rrf(lists, k = 60) {
  const scores = new Map();
  for (const list of lists) {
    list.forEach((doc, rank) => {
      const prev = scores.get(doc.id) || { doc, score: 0 };
      prev.score += 1 / (k + rank + 1); // earlier rank -> bigger contribution
      scores.set(doc.id, prev);
    });
  }
  return [...scores.values()].sort((a, b) => b.score - a.score).map(x => x.doc);
}

const dense = await db.search(queryVector, { topK: 20 });
const sparse = await bm25.search(queryText, { topK: 20 });
const fused = rrf([dense, sparse]).slice(0, 5);

Diversity: MMR

Plain top-k can return five near-duplicate chunks that all say the same thing — wasting your context budget. Maximal Marginal Relevance (MMR) balances relevance against novelty, preferring results that are relevant and different from what's already selected. Useful when your corpus is repetitive or a question has multiple facets.

Filtering

Combine similarity with metadata predicates so you only search what's allowed and relevant — this tenant, this date range, these permitted sections. Filtering is also your first line of defence for access control: never retrieve what the user isn't allowed to see. See Security.

Score thresholds and empty results

Sometimes the best match is still a poor match. Applying a minimum similarity threshold lets you detect "nothing relevant found" and respond honestly ("I don't have information on that") instead of forcing the model to answer from irrelevant context. An empty or low-confidence retrieval is valuable signal, not a failure to paper over.

Improving retrieval, in order of effort

  1. Fix chunking (biggest lever, see Chunking).
  2. Add hybrid search (dense + sparse).
  3. Add a reranker over a larger candidate set.
  4. Add query transformations (rewrite, multi-query, HyDE).
  5. Only then consider fine-tuning embeddings.
Bipin Singh
Written by Bipin Singh

Senior Full-Stack Engineer · AI & AWS. I build production RAG and LLM systems for enterprises.

Work with me