Hybrid search
Vector search understands meaning; keyword search understands exact words. Each fails where the other succeeds. Hybrid search runs both and combines the results, and it is often the single biggest quality improvement you can make to a search or RAG system.
Where vector search struggles
| Query | Why pure vector search may fail |
|---|---|
ERR_CERT_AUTHORITY_INVALID |
Rare codes and identifiers carry little "meaning" for an embedding model |
Invoice INV-2026-0042 |
Exact IDs must match exactly |
| "pgvector hnsw ef_search" | Product names and parameters, especially new ones |
| A person's surname | Names may be poorly represented |
Keyword search handles these easily. Meanwhile, keyword search fails on "my card got declined" vs "payment refused" — where vector search shines.
Dense and sparse vectors
- Dense vectors come from embedding models: a few hundred to a few thousand numbers, almost all non-zero. They capture meaning.
- Sparse vectors represent text by the words (terms) it contains: a huge vocabulary-sized list where almost every entry is zero. Classic keyword ranking like BM25 works this way, as do learned sparse models (for example, SPLADE) that also add related terms.
Many vector databases now store sparse and dense vectors side by side, or pair with a keyword engine.
Combining the results
You run both searches and then need one ranked list. Their scores aren't on the same scale — a cosine of 0.82 and a BM25 score of 14.3 can't be added directly.
Reciprocal Rank Fusion (RRF)
RRF ignores raw scores and uses rank positions instead:
RRF score(doc) = Σ over each result list of 1 / (k + rank of doc in that list)
(k is a constant, commonly 60)
Worked example with k = 60:
| Document | Rank in vector results | Rank in keyword results | RRF score |
|---|---|---|---|
| A | 1 | 3 | 1/61 + 1/63 ≈ 0.0323 |
| B | 2 | — | 1/62 ≈ 0.0161 |
| C | 5 | 1 | 1/65 + 1/61 ≈ 0.0318 |
A and C rank highest because both searches liked them. RRF is simple, robust and needs no tuning — a great default.
Weighted score fusion
Normalise each list's scores (for example, to 0–1) and combine them with a weight, such as 0.7 × vector + 0.3 × keyword. More tunable, but sensitive to how scores are distributed; tune it against an evaluation set.
Add a reranker
A reranker (cross-encoder) reads the query and each candidate together and scores relevance directly. It is much more accurate than comparing embeddings, but too slow to run over a whole collection. The standard pipeline:
hybrid search → top 50 candidates → reranker → top 5 → user or LLM
A strong default for search and RAG: dense + keyword search, fused with RRF, then reranked. Each stage fixes weaknesses of the previous one.
When hybrid is worth it
- Your users search for codes, IDs, names or jargon.
- Your content is technical, legal or product-heavy.
- Evaluation shows exact-match queries failing.
Measure before and after: build a small set of real queries with known good answers and compare recall with and without hybrid. The RAG Handbook goes deeper on retrieval strategies.