RAG Handbook retrieval-augmented generation Bipin Singh
Advanced RAG

Reranking

2 min readChapter 12 of 26By Bipin Singh

Reranking is a two-stage retrieval pattern: cast a wide net cheaply, then apply a slower, more accurate model to reorder the top candidates. It's one of the most reliable quality boosts you can add, because it fixes the exact place where vector search is weakest — fine-grained ranking.

Why a second stage?

Vector search compares the query and each document independently — each was embedded on its own, before the query existed. That's fast and scalable but coarse: it can rank a loosely related passage above a precisely relevant one. A reranker looks at the query and a candidate together and scores their actual relevance, catching what the first pass missed.

Bi-encoders vs cross-encoders

Bi-encoder (retrieval)Cross-encoder (reranking)
HowEmbeds query and doc separately, compares vectorsReads query + doc together, outputs a relevance score
SpeedVery fast; vectors precomputedSlow; runs the model per (query, doc) pair
AccuracyGoodHigher — sees the interaction between query and doc
RoleSearch millionsReorder the top tens
Key idea
Retrieve with a fast bi-encoder over the whole corpus, then rerank with a slow cross-encoder over the top ~20–50. You get the scale of the first and the precision of the second.

The two-stage pattern

// Stage 1: cheap, wide recall.
const candidates = await db.search(queryVector, { topK: 50 });

// Stage 2: expensive, precise reordering.
const scored = await reranker.score(
  queryText,
  candidates.map(c => c.text)
);
const reranked = candidates
  .map((c, i) => ({ ...c, score: scored[i] }))
  .sort((a, b) => b.score - a.score)
  .slice(0, 5); // keep only the best few for the prompt

Rerankers in practice

You can use a hosted reranking API (e.g. Cohere Rerank, Voyage) or an open cross-encoder (e.g. the bge-reranker family) you host yourself. Trade-offs mirror embeddings: hosted is easy but per-call and off-network; self-hosted is private and fixed-cost but you run it. LLMs can also rerank by listing candidates and asking for the most relevant, though dedicated rerankers are usually cheaper and faster.

What you gain

The trade-off

Reranking adds latency and compute proportional to the number of candidates. Tune how many you retrieve before reranking (recall) against how many you keep after (precision and cost). It's usually well worth it — often more impactful than swapping in a fancier embedding model.

Bipin Singh
Written by Bipin Singh

Senior Full-Stack Engineer · AI & AWS. I build production RAG and LLM systems for enterprises.

Work with me