Vector Database Handbook from zero to production Bipin Singh
Choosing & running

Choosing a vector database

2 min readChapter 18 of 25By Bipin Singh

Choosing a vector database is less about which product is "best" and more about which one fits your data, team and constraints. This chapter gives you the questions to ask and a simple evaluation process.

Start with these questions

Question Why it matters
How many vectors, now and in two years? Thousands, millions and billions call for different approaches
How many dimensions? Drives memory and cost
What query latency and throughput do you need? A chat feature at 10 queries/second differs from search at 5,000
How fresh must results be? Real-time updates vs nightly batch loads
How heavily will you filter? Multi-tenant SaaS and permission-aware search need strong filtering
Do users search for exact terms, codes or names? You probably need hybrid search
What do you already run? Reusing Postgres or Elasticsearch avoids a new system
Who will operate it? Managed service vs self-hosted changes your on-call load
Where must the data live? Residency, VPC-only or on-premises requirements narrow the options
What's the budget? Memory-heavy indexes at scale get expensive

A simple decision guide

Do you already run PostgreSQL, and expect < ~10M vectors?
  └─ yes → start with pgvector
Do you already run Elasticsearch/OpenSearch and need strong keyword search?
  └─ yes → use its vector search for hybrid
Is vector search core to your product, at large scale or high QPS,
or do you need advanced filtering / multi-tenancy?
  └─ yes → a dedicated vector database (managed if you don't want to operate it)
Is it an offline batch job or a custom search engine?
  └─ yes → a library such as FAISS or hnswlib
Is it a prototype, local tool or notebook?
  └─ yes → an embedded option, or even brute force in memory

These thresholds are rough. Real limits depend on hardware, dimensions, filters and tuning — which is why you test.

Key idea

The most common good answer is "start with what you already run, measure, and move to a dedicated system only when you hit a real limit." Migrating vectors later is straightforward if your ingestion can rebuild the index from source.

Running a fair evaluation

  1. Use your data. A sample of your real documents and real queries — public benchmarks won't include your filters or your content.
  2. Build a ground truth. Exact (brute-force) top-k for each test query, plus a small set of human-judged relevant results if you can.
  3. Test realistic queries. Include your filters, tenant sizes and hybrid queries, not just unfiltered vector search.
  4. Measure at matched recall. Compare latency and cost at the same recall target (e.g. recall@10 ≥ 0.95).
  5. Test writes too. Ingestion speed, update/delete behaviour and how quickly new data becomes searchable.
  6. Load test. p95/p99 latency at your expected peak throughput, not just a single query.
  7. Count the operational cost. Backups, upgrades, monitoring, scaling, security reviews — and who does them.

Red flags during evaluation

Tip

Write down your decision and the evidence — dataset, recall target, latency, cost. It's useful for your team, and it makes an excellent case study for interviews.

Bipin Singh
Written by Bipin Singh

Senior Full-Stack Engineer · AI & AWS. I build production search, RAG and AI systems on AWS and Postgres.

Work with me