Evaluating search quality
There are two different questions you can ask about vector search:
- Is the index accurate? Does the ANN search return the same results an exact search would? (Index quality.)
- Are the results relevant? Do the returned items actually answer the user's need? (Search quality.)
A perfect index can still return irrelevant results if the embeddings or chunking are poor. You need to measure both.
Index quality: recall against exact search
Run each test query with exact (brute-force) search and with your ANN index, and compare:
recall@10 = |ANN top 10 ∩ exact top 10| ÷ 10
Average across all test queries. This needs no human labels — just compute power — so it's easy to automate.
Search quality: relevance metrics
For these you need relevance labels: for each test query, which items are actually relevant (judged by people who know the content, or by a carefully checked LLM grader).
| Metric | Question it answers | Example |
|---|---|---|
| Recall@k | Of all relevant items, how many appear in the top k? | 3 relevant docs; 2 appear in top 5 → 0.67 |
| Precision@k | Of the top k results, how many are relevant? | 2 of top 5 relevant → 0.4 |
| Hit rate@k | Does at least one relevant item appear in the top k? | Yes → 1 |
| MRR (mean reciprocal rank) | How high is the first relevant result? | First relevant at rank 3 → 1/3 |
| nDCG@k | Are the most relevant results ranked highest? (Supports graded relevance) | Rewards putting the best answer first |
For RAG, recall@k and hit rate@k matter most — the LLM can only use what retrieval gives it. For search UIs, MRR and nDCG matter more, because users look at the top results first.
Building a test set
- Collect real queries — from logs, support tickets or user interviews. Aim for a few hundred, covering common and tricky cases.
- Label relevant results — have domain experts mark which documents answer each query. Start with 50 well-labelled queries rather than 500 sloppy ones.
- Include hard cases — exact codes and names, typos, multi-part questions, queries with no good answer.
- Freeze it — version the test set so results are comparable over time.
An LLM can draft candidate questions from your documents, and help grade relevance at scale — but have a human review a sample, and keep a core set that's fully human-labelled.
What to compare
Use the same test set to compare:
- Embedding models
- Chunk sizes and strategies
- Dense vs hybrid search
- With and without a reranker
- Index settings (recall vs exact search)
Change one thing at a time, and report results in a small table so decisions are traceable.
Online signals
Offline metrics tell you whether a change should help; online signals tell you whether it did:
- Click-through on top results
- Users rephrasing and searching again (a sign of failure)
- Thumbs up/down on answers in RAG apps
- "No results" or fallback rates
Measure index recall to make search fast without losing accuracy, and measure relevance to make sure the results are actually useful. Both are needed; neither replaces the other.