RAG Handbook retrieval-augmented generation Bipin Singh
The core pipeline

Augmentation & prompting

2 min readChapter 09 of 26By Bipin Singh

Augmentation is the step where retrieved passages become part of the prompt. It's easy to treat this as trivial string concatenation, but how you assemble the context strongly affects answer quality, faithfulness, and cost.

A solid prompt structure

A grounded RAG prompt generally contains: instructions on how to use the context, the retrieved context itself (clearly delimited and labelled), and the user's question. Being explicit about grounding is what keeps the model honest.

You are a helpful assistant. Answer the question using ONLY the context below.
If the answer is not contained in the context, say you don't know — do not guess.
Cite the source number(s) you used, like [1].

Context:
[1] (returns.pdf, p.3) The refund window is 30 days from delivery...
[2] (faq.md) Refunds are issued to the original payment method...

Question: How long do I have to request a refund?

Instruct the model to stay grounded

Two instructions do most of the work: use only the provided context, and say "I don't know" when the answer isn't there. Without the second, models fill gaps with plausible fabrications. With it, an unanswered question becomes an honest non-answer instead of a confident wrong one.

Tip
Give the model an explicit escape hatch. "If the context doesn't contain the answer, say you don't know" is the single most effective line for reducing hallucination in RAG.

Context budgeting

You have a finite token budget shared between instructions, retrieved context, conversation history, and room for the answer. More context is not always better — it costs money, adds latency, and can bury the relevant passage. Retrieve broadly, then trim to the few passages that earn their place (reranking helps here). Track token counts so you never silently overflow the window.

Order matters — "lost in the middle"

Models attend most reliably to the beginning and end of the context and least reliably to the middle. When you include several passages, place the most important ones at the edges rather than buried in the centre. If you rerank, put the top result first (or last), not in the middle of the pile.

Citations

Because every passage carries metadata, you can ask the model to cite the source of each claim. Citations make answers verifiable, build user trust, and are often a hard requirement in regulated settings. Number your context passages and instruct the model to reference those numbers; then map them back to real sources in your UI.

Conversation history and query rewriting

In a chat, follow-up questions are often incomplete: "what about the enterprise plan?" means nothing on its own. Before retrieving, rewrite the follow-up into a standalone question using the conversation so far — otherwise you embed an ambiguous fragment and retrieve poorly. This is covered in Query transformations.

Assembling context in code

function buildPrompt(question, hits) {
  const context = hits
    .map((h, i) => "[" + (i + 1) + "] (" + h.metadata.source + ") " + h.text)
    .join("\n\n");

  return [
    "Answer using ONLY the context below. If it's not there, say you don't know.",
    "Cite sources by number, e.g. [1].",
    "",
    "Context:",
    context,
    "",
    "Question: " + question,
  ].join("\n");
}
Bipin Singh
Written by Bipin Singh

Senior Full-Stack Engineer · AI & AWS. I build production RAG and LLM systems for enterprises.

Work with me