RAG Handbook retrieval-augmented generation Bipin Singh
Foundations

RAG vs fine-tuning vs long context

2 min readChapter 03 of 26By Bipin Singh

RAG is not the only way to make a model useful on information it wasn't trained on. The three main options solve overlapping but distinct problems, and mature systems often combine them.

The three approaches

ApproachWhat it changesBest forWeakness
Prompting / long contextWhat you put in the prompt each callSmall, known context that fits in the windowCost and latency scale with context; hard limits; "lost in the middle"
RAGWhat the model can look up at query timeLarge, private, or changing knowledge that must be citedRetrieval can miss; more moving parts to operate
Fine-tuningThe model's weights (behaviour, style, format)Teaching a skill, tone, or output shapeExpensive to update; knowledge goes stale; risk of forgetting

The key distinction: knowledge vs behaviour

This one framing resolves most confusion:

Key idea
Fine-tuning teaches the model how to act; RAG teaches it what to reference. Trying to fine-tune facts in is a common, costly mistake — the facts are frozen at training time and drift out of date.

RAG vs stuffing everything in a long context window

Modern models accept very large contexts, which tempts people to skip retrieval and paste the whole corpus in. That works until it doesn't:

Long context and RAG are complements: retrieval narrows millions of tokens down to the few thousand that matter, and a large window gives you room to include more of them plus room for reasoning.

A quick decision guide

Does the answer depend on private / changing / large knowledge?
  |-- No  -> plain prompting; add long context if the material is small & static
  |-- Yes -> use RAG
Do you also need a specific tone, format, or narrow skill?
  |-- Yes -> add fine-tuning ON TOP of RAG (behaviour), keep facts in RAG (knowledge)

In short: default to RAG for knowledge, layer fine-tuning for behaviour, and use the context window as the space where retrieved knowledge and model reasoning meet.

Bipin Singh
Written by Bipin Singh

Senior Full-Stack Engineer · AI & AWS. I build production RAG and LLM systems for enterprises.

Work with me