Augmentation & prompting
Augmentation is the step where retrieved passages become part of the prompt. It's easy to treat this as trivial string concatenation, but how you assemble the context strongly affects answer quality, faithfulness, and cost.
A solid prompt structure
A grounded RAG prompt generally contains: instructions on how to use the context, the retrieved context itself (clearly delimited and labelled), and the user's question. Being explicit about grounding is what keeps the model honest.
You are a helpful assistant. Answer the question using ONLY the context below.
If the answer is not contained in the context, say you don't know — do not guess.
Cite the source number(s) you used, like [1].
Context:
[1] (returns.pdf, p.3) The refund window is 30 days from delivery...
[2] (faq.md) Refunds are issued to the original payment method...
Question: How long do I have to request a refund?
Instruct the model to stay grounded
Two instructions do most of the work: use only the provided context, and say "I don't know" when the answer isn't there. Without the second, models fill gaps with plausible fabrications. With it, an unanswered question becomes an honest non-answer instead of a confident wrong one.
Context budgeting
You have a finite token budget shared between instructions, retrieved context, conversation history, and room for the answer. More context is not always better — it costs money, adds latency, and can bury the relevant passage. Retrieve broadly, then trim to the few passages that earn their place (reranking helps here). Track token counts so you never silently overflow the window.
Order matters — "lost in the middle"
Models attend most reliably to the beginning and end of the context and least reliably to the middle. When you include several passages, place the most important ones at the edges rather than buried in the centre. If you rerank, put the top result first (or last), not in the middle of the pile.
Citations
Because every passage carries metadata, you can ask the model to cite the source of each claim. Citations make answers verifiable, build user trust, and are often a hard requirement in regulated settings. Number your context passages and instruct the model to reference those numbers; then map them back to real sources in your UI.
Conversation history and query rewriting
In a chat, follow-up questions are often incomplete: "what about the enterprise plan?" means nothing on its own. Before retrieving, rewrite the follow-up into a standalone question using the conversation so far — otherwise you embed an ambiguous fragment and retrieve poorly. This is covered in Query transformations.
Assembling context in code
function buildPrompt(question, hits) {
const context = hits
.map((h, i) => "[" + (i + 1) + "] (" + h.metadata.source + ") " + h.text)
.join("\n\n");
return [
"Answer using ONLY the context below. If it's not there, say you don't know.",
"Cite sources by number, e.g. [1].",
"",
"Context:",
context,
"",
"Question: " + question,
].join("\n");
}