Vector Database Handbook from zero to production Bipin Singh
Choosing & running

Performance tuning

2 min readChapter 20 of 25By Bipin Singh

Tuning a vector database is a loop: measure, change one thing, measure again. Without measurement, "tuning" is guessing. This chapter gives you a repeatable method.

Set targets first

Agree on targets that come from the product:

The tuning loop

  1. Build a test set of a few hundred realistic queries, with exact top-k results computed by brute force.
  2. Measure the baseline — recall, p50/p95/p99 latency, QPS.
  3. Change one setting.
  4. Re-measure and record the result.
  5. Stop when you hit your targets with some headroom.

Levers, in the order to try them

Order Lever Effect
1 Query-time search width (efSearch, nprobe) Directly trades recall for latency; no rebuild
2 Result count and filters Asking for fewer results, or filtering smarter, reduces work
3 Payload indexes on filter fields Speeds up filtered search dramatically
4 Quantization with rescoring Less memory and faster comparisons, small accuracy cost
5 Build-time parameters (M, efConstruction, nlist) Better graph/clusters; requires rebuild
6 Hardware: more RAM, faster CPUs, more replicas More throughput and headroom
7 Fewer dimensions or a different embedding model Big change; re-embed and re-evaluate
Key idea

Always report latency at a recall level. "5 ms" means nothing on its own; "5 ms at recall@10 = 0.96" is a result.

Plotting the trade-off

Sweep the main knob and record each point:

efSearch recall@10 p95 latency
16 0.86 1.9 ms
32 0.92 2.6 ms
64 0.96 3.8 ms
128 0.98 6.5 ms
256 0.99 11.8 ms

(Illustrative numbers.) Pick the cheapest point that meets your target — here, efSearch = 64. Notice the diminishing returns: the last 1–2% of recall often costs as much latency as everything before it.

Common performance problems

Symptom Likely cause Fix
Latency spikes after bulk load Index still building or compacting Load first, build indexes after; schedule compaction
Filtered queries slow or low recall Filter field not indexed; very selective filter Add payload indexes; pre-filter + exact search for small subsets
Memory keeps growing Deletes/updates leaving tombstones Run compaction or optimise; rebuild periodically
Good recall in tests, poor in production Test queries unrealistic Sample real queries for the test set
Slow first queries after restart Index not yet loaded in memory (cold cache) Warm up with representative queries
Tip

Re-run your tuning test set whenever you change the embedding model, database version, quantization or index parameters. Treat it like a regression test suite.

Bipin Singh
Written by Bipin Singh

Senior Full-Stack Engineer · AI & AWS. I build production search, RAG and AI systems on AWS and Postgres.

Work with me