Foundations
RAG vs fine-tuning vs long context
RAG is not the only way to make a model useful on information it wasn't trained on. The three main options solve overlapping but distinct problems, and mature systems often combine them.
The three approaches
| Approach | What it changes | Best for | Weakness |
|---|---|---|---|
| Prompting / long context | What you put in the prompt each call | Small, known context that fits in the window | Cost and latency scale with context; hard limits; "lost in the middle" |
| RAG | What the model can look up at query time | Large, private, or changing knowledge that must be cited | Retrieval can miss; more moving parts to operate |
| Fine-tuning | The model's weights (behaviour, style, format) | Teaching a skill, tone, or output shape | Expensive to update; knowledge goes stale; risk of forgetting |
The key distinction: knowledge vs behaviour
This one framing resolves most confusion:
- Need the model to know something? That is a knowledge problem — use RAG (or long context for small amounts). Facts, documents, current data.
- Need the model to behave a certain way? That is a behaviour problem — use fine-tuning (or careful prompting). Style, format, tone, a narrow task done consistently.
Key idea
Fine-tuning teaches the model how to act; RAG teaches it what to reference. Trying to fine-tune facts in is a common, costly mistake — the facts are frozen at training time and drift out of date.
RAG vs stuffing everything in a long context window
Modern models accept very large contexts, which tempts people to skip retrieval and paste the whole corpus in. That works until it doesn't:
- Cost. You pay per input token on every call. Sending 200k tokens to answer a one-line question is wasteful when 2k relevant tokens would do.
- Latency. More input means slower responses.
- Attention dilution. Models attend less reliably to information buried in the middle of a huge context — the "lost in the middle" effect. A tight, relevant context often beats a giant one.
- Ceiling. Corpora are frequently larger than any context window. Retrieval scales; stuffing does not.
Long context and RAG are complements: retrieval narrows millions of tokens down to the few thousand that matter, and a large window gives you room to include more of them plus room for reasoning.
A quick decision guide
Does the answer depend on private / changing / large knowledge?
|-- No -> plain prompting; add long context if the material is small & static
|-- Yes -> use RAG
Do you also need a specific tone, format, or narrow skill?
|-- Yes -> add fine-tuning ON TOP of RAG (behaviour), keep facts in RAG (knowledge)
In short: default to RAG for knowledge, layer fine-tuning for behaviour, and use the context window as the space where retrieved knowledge and model reasoning meet.