RAG Handbook retrieval-augmented generation Bipin Singh
The core pipeline

Generation

2 min readChapter 10 of 26By Bipin Singh

Generation is where the LLM reads the assembled prompt and writes the answer. The retrieval work you've done sets the ceiling on quality; generation determines how much of that ceiling you actually reach — how faithful, clear, and useful the final answer is.

Choosing the generation model

Balance capability, latency, and cost against the task. A hard reasoning task over dense technical passages needs a stronger model; a simple lookup can use a smaller, faster, cheaper one. A common pattern is tiered: a small model for easy queries, escalating to a larger one when needed (routing can decide — see Routing).

Faithfulness: answering from the context

The defining risk of RAG generation is the model ignoring the retrieved context and answering from its own (possibly wrong or outdated) parametric memory. You reduce this with prompt discipline (see Augmentation) and verify it with evaluation:

Both are measurable; see Evaluation.

Watch out
A fluent answer is not a correct answer. RAG can produce beautifully written responses that are unsupported by the sources. Measure faithfulness explicitly — never assume it.

Handling "no good context"

When retrieval returns nothing relevant (or scores below threshold), the right behaviour is usually to say so, not to improvise. Decide the policy deliberately: refuse, ask a clarifying question, or fall back to a general answer clearly labelled as not sourced from your documents. Silent fabrication is the worst outcome.

Structured output

Many RAG applications need more than prose — JSON for downstream code, fields for a form, a list of citations. Ask for a specific schema and validate the result. Structured output also makes citations machine-checkable: return the answer and the exact source IDs it relied on, then verify programmatically that those IDs were in the retrieved set.

// Request a checkable, structured answer.
const schema = {
  answer: "string",
  sources: "array of source numbers actually used",
  answered: "boolean — false if context was insufficient",
};
// After generation: assert every returned source id was in `hits`,
// and if answered === false, surface an honest 'no answer' to the user.

Streaming and UX

Retrieval adds latency before the first token. Stream the answer as it's generated so the interface feels responsive, and show which sources are being used. Displaying citations inline lets users verify claims themselves — turning the model from an oracle into a research assistant they can check.

Post-generation checks

For higher-stakes uses, add a verification pass: a second model call (or rules) that checks whether the answer's claims are grounded in the context, flags unsupported statements, and can trigger a retry with different retrieval. This is the seed of the self-correcting architectures in Advanced architectures.

Bipin Singh
Written by Bipin Singh

Senior Full-Stack Engineer · AI & AWS. I build production RAG and LLM systems for enterprises.

Work with me