RAG Handbook retrieval-augmented generation Bipin Singh
Evaluation & production

Production & scaling

2 min readChapter 19 of 26By Bipin Singh

A demo that works on ten documents is a different animal from a system serving thousands of users over millions of documents. Production RAG is an engineering problem: latency budgets, cost control, and keeping knowledge current without downtime.

Latency

Each stage adds delay: query transformation, embedding, vector search, reranking, and generation. Generation usually dominates. Tactics:

Caching — at several layers

Cost control

The main cost drivers are LLM tokens (input + output) and, at scale, vector storage and embedding calls. Levers: retrieve less but better (reranking lets you send fewer, higher-quality chunks), pick right-sized models per query, cache aggressively, and compress or truncate embeddings (Matryoshka) to shrink storage. Measure cost per query and watch it like any other production metric.

Scaling ingestion

Indexing millions of documents is a batch data pipeline: parallelise parsing and embedding, batch embedding calls for throughput, handle failures and retries, and track progress so you can resume. Treat it like ETL, because it is.

Freshness & incremental indexing

Rebuilding the whole index on every change doesn't scale. Instead:

// Incremental update: only touch what changed.
for (const doc of sourceDocs) {
  const hash = sha256(doc.content);
  const existing = await index.getMeta(doc.id);
  if (!existing)              await indexDocument(doc);        // new
  else if (existing.hash !== hash) await reindexDocument(doc); // changed
  // unchanged -> skip
}
await index.deleteMissing(sourceDocs.map(d => d.id));          // removed

Reliability

External embedders and LLMs fail and rate-limit. Add timeouts, retries with backoff, and graceful fallbacks (a smaller model, a cached answer, or an honest "try again"). Decide what happens when the vector DB is briefly unavailable — degrade, don't crash.

Bipin Singh
Written by Bipin Singh

Senior Full-Stack Engineer · AI & AWS. I build production RAG and LLM systems for enterprises.

Work with me