Choosing & running
Choosing a vector database
Choosing a vector database is less about which product is "best" and more about which one fits your data, team and constraints. This chapter gives you the questions to ask and a simple evaluation process.
Start with these questions
| Question | Why it matters |
|---|---|
| How many vectors, now and in two years? | Thousands, millions and billions call for different approaches |
| How many dimensions? | Drives memory and cost |
| What query latency and throughput do you need? | A chat feature at 10 queries/second differs from search at 5,000 |
| How fresh must results be? | Real-time updates vs nightly batch loads |
| How heavily will you filter? | Multi-tenant SaaS and permission-aware search need strong filtering |
| Do users search for exact terms, codes or names? | You probably need hybrid search |
| What do you already run? | Reusing Postgres or Elasticsearch avoids a new system |
| Who will operate it? | Managed service vs self-hosted changes your on-call load |
| Where must the data live? | Residency, VPC-only or on-premises requirements narrow the options |
| What's the budget? | Memory-heavy indexes at scale get expensive |
A simple decision guide
Do you already run PostgreSQL, and expect < ~10M vectors?
└─ yes → start with pgvector
Do you already run Elasticsearch/OpenSearch and need strong keyword search?
└─ yes → use its vector search for hybrid
Is vector search core to your product, at large scale or high QPS,
or do you need advanced filtering / multi-tenancy?
└─ yes → a dedicated vector database (managed if you don't want to operate it)
Is it an offline batch job or a custom search engine?
└─ yes → a library such as FAISS or hnswlib
Is it a prototype, local tool or notebook?
└─ yes → an embedded option, or even brute force in memory
These thresholds are rough. Real limits depend on hardware, dimensions, filters and tuning — which is why you test.
Key idea
The most common good answer is "start with what you already run, measure, and move to a dedicated system only when you hit a real limit." Migrating vectors later is straightforward if your ingestion can rebuild the index from source.
Running a fair evaluation
- Use your data. A sample of your real documents and real queries — public benchmarks won't include your filters or your content.
- Build a ground truth. Exact (brute-force) top-k for each test query, plus a small set of human-judged relevant results if you can.
- Test realistic queries. Include your filters, tenant sizes and hybrid queries, not just unfiltered vector search.
- Measure at matched recall. Compare latency and cost at the same recall target (e.g. recall@10 ≥ 0.95).
- Test writes too. Ingestion speed, update/delete behaviour and how quickly new data becomes searchable.
- Load test. p95/p99 latency at your expected peak throughput, not just a single query.
- Count the operational cost. Backups, upgrades, monitoring, scaling, security reviews — and who does them.
Red flags during evaluation
- Recall drops sharply with your real filters.
- Latency is great at low load but degrades badly near your expected peak.
- Deletes or updates aren't reflected for a long time, or degrade quality over time.
- No clear path for backups, disaster recovery or embedding-model migration.
- Can't be deployed where your data must live.
Tip
Write down your decision and the evidence — dataset, recall target, latency, cost. It's useful for your team, and it makes an excellent case study for interviews.