Running it in production
A vector database in production is a database like any other: it needs backups, monitoring, security and a plan for change. Vector search also brings a few operational concerns of its own — especially embedding model changes.
Backups and recovery
- Snapshots: use the database's snapshot or backup feature, and test restoring from it.
- Rebuild from source: because the vector database is a search index, keep the ability to regenerate it from your source data. This is your ultimate backup — just remember that re-embedding a large corpus takes time and costs money, so snapshots are still valuable for fast recovery.
- Know your recovery time: how long would a full restore or rebuild take? Is that acceptable?
Monitoring
| What to watch | Why |
|---|---|
| Query latency (p50/p95/p99) | User experience; early sign of overload |
| Queries per second and error rate | Capacity and failures |
| Memory and disk usage | In-memory indexes fail badly when memory runs out |
| Index build / compaction status | Background work affects latency and freshness |
| Ingestion lag | How stale search results are |
| Recall on a canary test set | Silent quality degradation |
| Empty-result rate | Broken filters or missing data |
| Embedding API latency and errors | Queries can't run without query embeddings |
Run a small set of known queries every few minutes and alert if expected results disappear. It catches broken ingestion, bad deployments and filter bugs before users do.
Security
- Authentication and network isolation — no public, unauthenticated endpoints; use private networking where possible.
- Encryption in transit and at rest.
- Permission-aware retrieval — enforce tenant and access filters inside search (see Metadata filtering).
- Treat embeddings as sensitive. Embeddings are derived from your data, and research has shown that original text can sometimes be partially reconstructed from them. Protect them like the source data.
- Audit logging for access to sensitive collections.
Changing embedding models safely
Embedding models improve, and providers deprecate old ones. Plan for migration from day one:
- Store the model name and version with every vector.
- Create a new collection or named vector for the new model.
- Re-embed in the background, rate-limited.
- Compare quality on your evaluation set.
- Switch traffic gradually; keep the old index until you're confident.
- Delete the old index.
Upgrades and schema changes
Test database upgrades on a copy with production-like data and your recall test set. Some upgrades require index rebuilds; schedule them for low-traffic periods, or use a blue–green setup: build the new index alongside the old and switch over.
Cost control
- Right-size memory after measuring real usage.
- Use quantization where accuracy allows.
- Avoid re-embedding unchanged content (hash chunks).
- Cache embeddings for frequent queries.
- Review idle collections and tenants.
Production readiness checklist
- Backups configured and a restore tested
- Index can be rebuilt from source with a documented procedure
- Latency, errors, memory and ingestion lag monitored with alerts
- Canary recall checks running
- Authentication, network isolation and encryption in place
- Tenant and permission filters enforced centrally and tested
- Embedding model version stored with vectors; migration plan written
- Load-tested at expected peak with realistic filters
- Runbook for common incidents (memory pressure, slow queries, stale data)
Treat the vector database as a rebuildable search index with production-grade operations around it. The teams that struggle are the ones who treated it as a demo that happened to go live.