RAG Handbook retrieval-augmented generation Bipin Singh
Foundations

How RAG works

2 min readChapter 02 of 26By Bipin Singh

RAG has an offline half and an online half. The offline half prepares your knowledge; the online half answers questions with it. Here is the whole flow.

Phase 1 — Indexing (offline)

This runs ahead of time, in a batch or on a schedule, whenever your source content changes.

  1. Load. Pull raw content from its source: files, a database, a website, an API.
  2. Parse & clean. Extract usable text from PDFs, HTML, Word docs, and so on; strip boilerplate.
  3. Chunk. Split long documents into passages small enough to retrieve precisely but large enough to carry meaning.
  4. Embed. Run each chunk through an embedding model to get a vector — a list of numbers that captures its meaning.
  5. Store. Save the vectors, the original text, and any metadata in a vector database.

Phase 2 — Retrieval & generation (online)

This runs per request, in the hot path, and must be fast.

  1. Embed the query. Convert the user's question into a vector using the same embedding model you used for the chunks.
  2. Search. Find the chunks whose vectors are closest to the query vector — the nearest neighbours in meaning.
  3. Augment. Assemble a prompt that contains the question plus the retrieved chunks as context.
  4. Generate. Send the prompt to the LLM, which writes an answer grounded in the supplied context.
  5. Return. Deliver the answer, ideally with citations back to the source chunks.
Tip
Use the same embedding model for indexing and for queries. Vectors from two different models live in different spaces and cannot be compared meaningfully.

A minimal end-to-end sketch

Stripped of frameworks, the online path looks like this:

async function answer(question, db, llm, embedder) {
  // 1. Embed the question with the same model used for indexing.
  const queryVector = await embedder.embed(question);

  // 2. Retrieve the most similar chunks.
  const hits = await db.search(queryVector, { topK: 5 });

  // 3. Build a grounded prompt.
  const context = hits.map((h, i) => "[" + (i + 1) + "] " + h.text).join("\n\n");
  const prompt =
    "Answer the question using ONLY the context below. " +
    "If the answer is not in the context, say you do not know.\n\n" +
    "Context:\n" + context + "\n\n" +
    "Question: " + question;

  // 4. Generate.
  return llm.complete(prompt);
}

The data flow at a glance

Reading left to right: content becomes vectors once; questions become vectors every time and are matched against them.

INDEXING (offline, batch)
  docs -> parse -> chunk -> embed -> [vector DB]

QUERY (online, per request)
  question -> embed --------> [vector DB] --search--> top-k chunks
                                                          |
                                    question + chunks --> LLM --> answer (+ citations)

Where quality is won or lost

Newcomers assume the LLM is the hard part. In practice, retrieval quality dominates. If the right passage never makes it into the prompt, no model — however capable — can answer correctly. That is why most of these docs are about the parts around the LLM: chunking, embeddings, search, and reranking. Garbage retrieved, garbage generated.

Bipin Singh
Written by Bipin Singh

Senior Full-Stack Engineer · AI & AWS. I build production RAG and LLM systems for enterprises.

Work with me