What is RAG?
Retrieval-Augmented Generation (RAG) is a technique that gives a language model access to information it was never trained on. Instead of relying only on what the model memorised during training, you retrieve relevant text at question time and place it in the prompt, so the model can generate an answer grounded in that text.
The problem it solves
A large language model is a snapshot. It knows what was in its training data up to a cutoff date, and nothing after. It has never seen your company's wiki, last night's support tickets, or the PDF a user just uploaded. Ask it about any of those and one of three things happens:
- It refuses — "I don't have information about that."
- It guesses — producing fluent, confident, wrong answers. This is hallucination.
- It's stale — answering from an outdated version of the world.
RAG addresses all three by moving knowledge out of the model's weights and into a searchable store you control. The model becomes a reasoning engine over text you supply, rather than a fallible memory bank.
Why not just retrain the model?
You could bake new knowledge into a model by fine-tuning or retraining, but that is slow, expensive, and hard to keep current. Knowledge changes daily; model training does not. RAG lets you update what the system knows by updating a database — no GPUs, no retraining, no waiting. You can add a document and have it answerable seconds later, and you can delete a document and have it forgotten just as fast (which matters for privacy and compliance).
When RAG is the right tool
Reach for RAG when your answers depend on a corpus that is large, private, or fast-changing:
- Question answering over internal docs, policies, or knowledge bases.
- Customer support grounded in product documentation.
- Search assistants over research papers, contracts, or codebases.
- Chatbots that must cite sources and stay current.
- Any case where "I made that up" is unacceptable and traceability matters.
When RAG is not the answer
- Pure skill, not knowledge. If you want the model to write in a certain style or follow a rigid format, that is a job for prompting or fine-tuning, not retrieval.
- Tiny, static context. If everything the model needs fits comfortably in the prompt every time, just put it there. RAG adds moving parts you don't need.
- Reasoning over the whole corpus at once. "Summarise all 10,000 documents" is not a retrieval problem — retrieval fetches a handful of passages, not the entire library.
The two halves of every RAG system
Every RAG system, however fancy, is built from two phases. Understanding this split is the foundation for everything else in these docs.
- Indexing (offline): load your documents, split them into chunks, convert each chunk into a vector (an embedding), and store the vectors in a database. You do this ahead of time and repeat it whenever content changes.
- Retrieval & generation (online): when a question arrives, embed the question, find the most similar chunks in the database, stuff them into a prompt, and ask the model to answer using them.
The next page walks through both phases end to end.