RAG Handbook retrieval-augmented generation Bipin Singh
Foundations

What is RAG?

3 min readChapter 01 of 26By Bipin Singh

Retrieval-Augmented Generation (RAG) is a technique that gives a language model access to information it was never trained on. Instead of relying only on what the model memorised during training, you retrieve relevant text at question time and place it in the prompt, so the model can generate an answer grounded in that text.

The problem it solves

A large language model is a snapshot. It knows what was in its training data up to a cutoff date, and nothing after. It has never seen your company's wiki, last night's support tickets, or the PDF a user just uploaded. Ask it about any of those and one of three things happens:

RAG addresses all three by moving knowledge out of the model's weights and into a searchable store you control. The model becomes a reasoning engine over text you supply, rather than a fallible memory bank.

Key idea
The one-sentence mental model: RAG is an open-book exam. The model is the student; retrieval hands it the right page of the textbook right before it answers.

Why not just retrain the model?

You could bake new knowledge into a model by fine-tuning or retraining, but that is slow, expensive, and hard to keep current. Knowledge changes daily; model training does not. RAG lets you update what the system knows by updating a database — no GPUs, no retraining, no waiting. You can add a document and have it answerable seconds later, and you can delete a document and have it forgotten just as fast (which matters for privacy and compliance).

When RAG is the right tool

Reach for RAG when your answers depend on a corpus that is large, private, or fast-changing:

When RAG is not the answer

The two halves of every RAG system

Every RAG system, however fancy, is built from two phases. Understanding this split is the foundation for everything else in these docs.

  1. Indexing (offline): load your documents, split them into chunks, convert each chunk into a vector (an embedding), and store the vectors in a database. You do this ahead of time and repeat it whenever content changes.
  2. Retrieval & generation (online): when a question arrives, embed the question, find the most similar chunks in the database, stuff them into a prompt, and ask the model to answer using them.

The next page walks through both phases end to end.

Bipin Singh
Written by Bipin Singh

Senior Full-Stack Engineer · AI & AWS. I build production RAG and LLM systems for enterprises.

Work with me