RAG Handbook retrieval-augmented generation Bipin Singh
The core pipeline

Vector databases

2 min readChapter 07 of 26By Bipin Singh

A vector database stores your embeddings and finds the nearest ones to a query vector — fast, even across millions of entries. The magic is approximate nearest-neighbour search: giving up a sliver of accuracy to make search orders of magnitude quicker.

What they hold

Each record typically bundles three things: the vector (for similarity search), the original text (to feed the LLM), and metadata (for filtering, citations, and access control). Keeping text alongside the vector saves a second lookup at query time.

Why approximate search?

Comparing a query against every vector (exact / brute-force search) is accurate but scales linearly — fine for thousands, painful for millions. Approximate nearest neighbour (ANN) algorithms build an index that finds almost the closest vectors in a fraction of the time. You trade a controllable amount of recall for a large speed gain.

Index algorithms you'll meet

IndexHow it worksTrade-off
FlatBrute-force, compares everythingPerfectly accurate, slow at scale — good for small sets
IVFClusters vectors; searches only nearby clustersFast; may miss neighbours near cluster edges
HNSWNavigable small-world graph you traverse toward the queryExcellent speed/recall; higher memory use — the common default
PQ / compressionQuantises vectors to use less memoryBig memory savings; small accuracy loss; often combined with IVF/HNSW
Tip
HNSW exposes tuning knobs (like how many neighbours to explore). More exploration = higher recall but slower queries. Tune it against your recall target, not by guesswork.

Metadata filtering

Real queries are rarely "search everything." You filter: only this tenant's documents, only docs newer than a date, only sections the user is allowed to see. Good vector DBs let you combine a similarity search with metadata predicates in one query. Pre-filtering (restrict candidates before the vector search) preserves result quality; naive post-filtering (search then drop) can leave you with too few results. Prefer engines that filter efficiently during search.

const hits = await db.search(queryVector, {
  topK: 5,
  filter: {
    tenantId: "acme",
    accessRoles: { in: userRoles },
    createdAt: { gte: "2024-01-01" },
  },
});

The landscape of options

Key idea
If you already run PostgreSQL, start with pgvector. Keeping vectors next to your relational data — one database, one backup, real transactions and joins — removes a whole class of operational headaches. Reach for a dedicated vector DB when scale or specialised features demand it.

Choosing one

Judge candidates on: scale (how many vectors, how many queries per second), filtering power, hybrid-search support, managed vs self-hosted, cost model, and how well it fits your existing stack. For most teams starting out, "the database we already operate, plus a vector extension" beats adding new infrastructure.

Bipin Singh
Written by Bipin Singh

Senior Full-Stack Engineer · AI & AWS. I build production RAG and LLM systems for enterprises.

Work with me