Choosing & running
Performance tuning
Tuning a vector database is a loop: measure, change one thing, measure again. Without measurement, "tuning" is guessing. This chapter gives you a repeatable method.
Set targets first
Agree on targets that come from the product:
- Quality: recall@k against exact search (e.g. recall@10 ≥ 0.95), and ideally end-to-end relevance on a labelled set.
- Latency: p95 (and p99) query latency at expected load (e.g. p95 < 50 ms).
- Throughput: queries per second at peak.
- Cost: memory and compute budget.
The tuning loop
- Build a test set of a few hundred realistic queries, with exact top-k results computed by brute force.
- Measure the baseline — recall, p50/p95/p99 latency, QPS.
- Change one setting.
- Re-measure and record the result.
- Stop when you hit your targets with some headroom.
Levers, in the order to try them
| Order | Lever | Effect |
|---|---|---|
| 1 | Query-time search width (efSearch, nprobe) |
Directly trades recall for latency; no rebuild |
| 2 | Result count and filters | Asking for fewer results, or filtering smarter, reduces work |
| 3 | Payload indexes on filter fields | Speeds up filtered search dramatically |
| 4 | Quantization with rescoring | Less memory and faster comparisons, small accuracy cost |
| 5 | Build-time parameters (M, efConstruction, nlist) |
Better graph/clusters; requires rebuild |
| 6 | Hardware: more RAM, faster CPUs, more replicas | More throughput and headroom |
| 7 | Fewer dimensions or a different embedding model | Big change; re-embed and re-evaluate |
Key idea
Always report latency at a recall level. "5 ms" means nothing on its own; "5 ms at recall@10 = 0.96" is a result.
Plotting the trade-off
Sweep the main knob and record each point:
| efSearch | recall@10 | p95 latency |
|---|---|---|
| 16 | 0.86 | 1.9 ms |
| 32 | 0.92 | 2.6 ms |
| 64 | 0.96 | 3.8 ms |
| 128 | 0.98 | 6.5 ms |
| 256 | 0.99 | 11.8 ms |
(Illustrative numbers.) Pick the cheapest point that meets your target — here, efSearch = 64. Notice the diminishing returns: the last 1–2% of recall often costs as much latency as everything before it.
Common performance problems
| Symptom | Likely cause | Fix |
|---|---|---|
| Latency spikes after bulk load | Index still building or compacting | Load first, build indexes after; schedule compaction |
| Filtered queries slow or low recall | Filter field not indexed; very selective filter | Add payload indexes; pre-filter + exact search for small subsets |
| Memory keeps growing | Deletes/updates leaving tombstones | Run compaction or optimise; rebuild periodically |
| Good recall in tests, poor in production | Test queries unrealistic | Sample real queries for the test set |
| Slow first queries after restart | Index not yet loaded in memory (cold cache) | Warm up with representative queries |
Tip
Re-run your tuning test set whenever you change the embedding model, database version, quantization or index parameters. Treat it like a regression test suite.