Production & scaling
A demo that works on ten documents is a different animal from a system serving thousands of users over millions of documents. Production RAG is an engineering problem: latency budgets, cost control, and keeping knowledge current without downtime.
Latency
Each stage adds delay: query transformation, embedding, vector search, reranking, and generation. Generation usually dominates. Tactics:
- Stream the answer so the user sees tokens immediately.
- Run independent steps in parallel (e.g. dense and sparse retrieval together).
- Use a smaller/faster model where quality allows; route hard queries to bigger ones.
- Keep top-k and reranking candidates as small as accuracy permits.
- Co-locate services to cut network hops between embedder, DB, and LLM.
Caching — at several layers
- Embedding cache: identical text embeds to the same vector — cache to avoid recomputing (helps repeated queries and re-indexing unchanged content).
- Retrieval cache: cache results for common or repeated queries.
- Semantic cache: if a new query is very similar to a past one, reuse the cached answer — big savings for FAQ-style traffic, but set the similarity threshold carefully to avoid serving a stale/wrong answer.
- Prompt caching: when supported, cache the static portion of prompts (system instructions, reused context) to cut input cost and latency.
Cost control
The main cost drivers are LLM tokens (input + output) and, at scale, vector storage and embedding calls. Levers: retrieve less but better (reranking lets you send fewer, higher-quality chunks), pick right-sized models per query, cache aggressively, and compress or truncate embeddings (Matryoshka) to shrink storage. Measure cost per query and watch it like any other production metric.
Scaling ingestion
Indexing millions of documents is a batch data pipeline: parallelise parsing and embedding, batch embedding calls for throughput, handle failures and retries, and track progress so you can resume. Treat it like ETL, because it is.
Freshness & incremental indexing
Rebuilding the whole index on every change doesn't scale. Instead:
- Give each source a stable ID and a content hash; re-embed only what changed.
- Delete removed content promptly — for correctness and for privacy/right-to-be-forgotten.
- Decide your freshness SLA: real-time updates cost more than a nightly batch. Match it to the use case.
// Incremental update: only touch what changed.
for (const doc of sourceDocs) {
const hash = sha256(doc.content);
const existing = await index.getMeta(doc.id);
if (!existing) await indexDocument(doc); // new
else if (existing.hash !== hash) await reindexDocument(doc); // changed
// unchanged -> skip
}
await index.deleteMissing(sourceDocs.map(d => d.id)); // removed
Reliability
External embedders and LLMs fail and rate-limit. Add timeouts, retries with backoff, and graceful fallbacks (a smaller model, a cached answer, or an honest "try again"). Decide what happens when the vector DB is briefly unavailable — degrade, don't crash.