Vector Database Handbook from zero to production Bipin Singh
Choosing & running

Running it in production

2 min readChapter 22 of 25By Bipin Singh

A vector database in production is a database like any other: it needs backups, monitoring, security and a plan for change. Vector search also brings a few operational concerns of its own — especially embedding model changes.

Backups and recovery

Monitoring

What to watch Why
Query latency (p50/p95/p99) User experience; early sign of overload
Queries per second and error rate Capacity and failures
Memory and disk usage In-memory indexes fail badly when memory runs out
Index build / compaction status Background work affects latency and freshness
Ingestion lag How stale search results are
Recall on a canary test set Silent quality degradation
Empty-result rate Broken filters or missing data
Embedding API latency and errors Queries can't run without query embeddings
Tip

Run a small set of known queries every few minutes and alert if expected results disappear. It catches broken ingestion, bad deployments and filter bugs before users do.

Security

Changing embedding models safely

Embedding models improve, and providers deprecate old ones. Plan for migration from day one:

  1. Store the model name and version with every vector.
  2. Create a new collection or named vector for the new model.
  3. Re-embed in the background, rate-limited.
  4. Compare quality on your evaluation set.
  5. Switch traffic gradually; keep the old index until you're confident.
  6. Delete the old index.

Upgrades and schema changes

Test database upgrades on a copy with production-like data and your recall test set. Some upgrades require index rebuilds; schedule them for low-traffic periods, or use a blue–green setup: build the new index alongside the old and switch over.

Cost control

Production readiness checklist

Key idea

Treat the vector database as a rebuildable search index with production-grade operations around it. The teams that struggle are the ones who treated it as a demo that happened to go live.

Bipin Singh
Written by Bipin Singh

Senior Full-Stack Engineer · AI & AWS. I build production search, RAG and AI systems on AWS and Postgres.

Work with me