Generation
Generation is where the LLM reads the assembled prompt and writes the answer. The retrieval work you've done sets the ceiling on quality; generation determines how much of that ceiling you actually reach — how faithful, clear, and useful the final answer is.
Choosing the generation model
Balance capability, latency, and cost against the task. A hard reasoning task over dense technical passages needs a stronger model; a simple lookup can use a smaller, faster, cheaper one. A common pattern is tiered: a small model for easy queries, escalating to a larger one when needed (routing can decide — see Routing).
Faithfulness: answering from the context
The defining risk of RAG generation is the model ignoring the retrieved context and answering from its own (possibly wrong or outdated) parametric memory. You reduce this with prompt discipline (see Augmentation) and verify it with evaluation:
- Faithfulness — is every claim in the answer supported by the retrieved context?
- Answer relevancy — does the answer actually address the question?
Both are measurable; see Evaluation.
Handling "no good context"
When retrieval returns nothing relevant (or scores below threshold), the right behaviour is usually to say so, not to improvise. Decide the policy deliberately: refuse, ask a clarifying question, or fall back to a general answer clearly labelled as not sourced from your documents. Silent fabrication is the worst outcome.
Structured output
Many RAG applications need more than prose — JSON for downstream code, fields for a form, a list of citations. Ask for a specific schema and validate the result. Structured output also makes citations machine-checkable: return the answer and the exact source IDs it relied on, then verify programmatically that those IDs were in the retrieved set.
// Request a checkable, structured answer.
const schema = {
answer: "string",
sources: "array of source numbers actually used",
answered: "boolean — false if context was insufficient",
};
// After generation: assert every returned source id was in `hits`,
// and if answered === false, surface an honest 'no answer' to the user.
Streaming and UX
Retrieval adds latency before the first token. Stream the answer as it's generated so the interface feels responsive, and show which sources are being used. Displaying citations inline lets users verify claims themselves — turning the model from an oracle into a research assistant they can check.
Post-generation checks
For higher-stakes uses, add a verification pass: a second model call (or rules) that checks whether the answer's claims are grounded in the context, flags unsupported statements, and can trigger a retry with different retrieval. This is the seed of the self-correcting architectures in Advanced architectures.