Vector databases
A vector database stores your embeddings and finds the nearest ones to a query vector — fast, even across millions of entries. The magic is approximate nearest-neighbour search: giving up a sliver of accuracy to make search orders of magnitude quicker.
What they hold
Each record typically bundles three things: the vector (for similarity search), the original text (to feed the LLM), and metadata (for filtering, citations, and access control). Keeping text alongside the vector saves a second lookup at query time.
Why approximate search?
Comparing a query against every vector (exact / brute-force search) is accurate but scales linearly — fine for thousands, painful for millions. Approximate nearest neighbour (ANN) algorithms build an index that finds almost the closest vectors in a fraction of the time. You trade a controllable amount of recall for a large speed gain.
Index algorithms you'll meet
| Index | How it works | Trade-off |
|---|---|---|
| Flat | Brute-force, compares everything | Perfectly accurate, slow at scale — good for small sets |
| IVF | Clusters vectors; searches only nearby clusters | Fast; may miss neighbours near cluster edges |
| HNSW | Navigable small-world graph you traverse toward the query | Excellent speed/recall; higher memory use — the common default |
| PQ / compression | Quantises vectors to use less memory | Big memory savings; small accuracy loss; often combined with IVF/HNSW |
Metadata filtering
Real queries are rarely "search everything." You filter: only this tenant's documents, only docs newer than a date, only sections the user is allowed to see. Good vector DBs let you combine a similarity search with metadata predicates in one query. Pre-filtering (restrict candidates before the vector search) preserves result quality; naive post-filtering (search then drop) can leave you with too few results. Prefer engines that filter efficiently during search.
const hits = await db.search(queryVector, {
topK: 5,
filter: {
tenantId: "acme",
accessRoles: { in: userRoles },
createdAt: { gte: "2024-01-01" },
},
});
The landscape of options
- Dedicated vector DBs: Pinecone, Weaviate, Qdrant, Milvus — built for vectors, rich filtering, scale.
- Embedded / local: Chroma, FAISS, LanceDB — great for prototyping and smaller apps.
- Extensions to databases you already run: pgvector (PostgreSQL), plus vector support in Elasticsearch, OpenSearch, Redis, MongoDB.
Choosing one
Judge candidates on: scale (how many vectors, how many queries per second), filtering power, hybrid-search support, managed vs self-hosted, cost model, and how well it fits your existing stack. For most teams starting out, "the database we already operate, plus a vector extension" beats adding new infrastructure.