Reranking
Reranking is a two-stage retrieval pattern: cast a wide net cheaply, then apply a slower, more accurate model to reorder the top candidates. It's one of the most reliable quality boosts you can add, because it fixes the exact place where vector search is weakest — fine-grained ranking.
Why a second stage?
Vector search compares the query and each document independently — each was embedded on its own, before the query existed. That's fast and scalable but coarse: it can rank a loosely related passage above a precisely relevant one. A reranker looks at the query and a candidate together and scores their actual relevance, catching what the first pass missed.
Bi-encoders vs cross-encoders
| Bi-encoder (retrieval) | Cross-encoder (reranking) | |
|---|---|---|
| How | Embeds query and doc separately, compares vectors | Reads query + doc together, outputs a relevance score |
| Speed | Very fast; vectors precomputed | Slow; runs the model per (query, doc) pair |
| Accuracy | Good | Higher — sees the interaction between query and doc |
| Role | Search millions | Reorder the top tens |
The two-stage pattern
// Stage 1: cheap, wide recall.
const candidates = await db.search(queryVector, { topK: 50 });
// Stage 2: expensive, precise reordering.
const scored = await reranker.score(
queryText,
candidates.map(c => c.text)
);
const reranked = candidates
.map((c, i) => ({ ...c, score: scored[i] }))
.sort((a, b) => b.score - a.score)
.slice(0, 5); // keep only the best few for the prompt
Rerankers in practice
You can use a hosted reranking API (e.g. Cohere Rerank, Voyage) or an open cross-encoder (e.g. the bge-reranker family) you host yourself. Trade-offs mirror embeddings: hosted is easy but per-call and off-network; self-hosted is private and fixed-cost but you run it. LLMs can also rerank by listing candidates and asking for the most relevant, though dedicated rerankers are usually cheaper and faster.
What you gain
- Higher precision in the top few results — the ones that actually reach the LLM.
- A smaller, cleaner context: rerank 50 down to 5, cutting noise and cost.
- A natural relevance threshold — drop candidates below a reranker score to detect "nothing relevant."
The trade-off
Reranking adds latency and compute proportional to the number of candidates. Tune how many you retrieve before reranking (recall) against how many you keep after (precision and cost). It's usually well worth it — often more impactful than swapping in a fancier embedding model.