Vector Database Handbook from zero to production Bipin Singh
Choosing & running

Evaluating search quality

2 min readChapter 21 of 25By Bipin Singh

There are two different questions you can ask about vector search:

  1. Is the index accurate? Does the ANN search return the same results an exact search would? (Index quality.)
  2. Are the results relevant? Do the returned items actually answer the user's need? (Search quality.)

A perfect index can still return irrelevant results if the embeddings or chunking are poor. You need to measure both.

Run each test query with exact (brute-force) search and with your ANN index, and compare:

recall@10 = |ANN top 10 ∩ exact top 10| ÷ 10

Average across all test queries. This needs no human labels — just compute power — so it's easy to automate.

Search quality: relevance metrics

For these you need relevance labels: for each test query, which items are actually relevant (judged by people who know the content, or by a carefully checked LLM grader).

Metric Question it answers Example
Recall@k Of all relevant items, how many appear in the top k? 3 relevant docs; 2 appear in top 5 → 0.67
Precision@k Of the top k results, how many are relevant? 2 of top 5 relevant → 0.4
Hit rate@k Does at least one relevant item appear in the top k? Yes → 1
MRR (mean reciprocal rank) How high is the first relevant result? First relevant at rank 3 → 1/3
nDCG@k Are the most relevant results ranked highest? (Supports graded relevance) Rewards putting the best answer first

For RAG, recall@k and hit rate@k matter most — the LLM can only use what retrieval gives it. For search UIs, MRR and nDCG matter more, because users look at the top results first.

Building a test set

  1. Collect real queries — from logs, support tickets or user interviews. Aim for a few hundred, covering common and tricky cases.
  2. Label relevant results — have domain experts mark which documents answer each query. Start with 50 well-labelled queries rather than 500 sloppy ones.
  3. Include hard cases — exact codes and names, typos, multi-part questions, queries with no good answer.
  4. Freeze it — version the test set so results are comparable over time.
Tip

An LLM can draft candidate questions from your documents, and help grade relevance at scale — but have a human review a sample, and keep a core set that's fully human-labelled.

What to compare

Use the same test set to compare:

Change one thing at a time, and report results in a small table so decisions are traceable.

Online signals

Offline metrics tell you whether a change should help; online signals tell you whether it did:

Key idea

Measure index recall to make search fast without losing accuracy, and measure relevance to make sure the results are actually useful. Both are needed; neither replaces the other.

Bipin Singh
Written by Bipin Singh

Senior Full-Stack Engineer · AI & AWS. I build production search, RAG and AI systems on AWS and Postgres.

Work with me