RAG Handbook retrieval-augmented generation Bipin Singh
Building & reference

Reference implementation

2 min readChapter 23 of 26By Bipin Singh

Here is a complete, deliberately minimal RAG pipeline in TypeScript — no framework, so every step is visible. It's a teaching skeleton, not production code, but it maps one-to-one to the concepts in these docs. Swap the placeholder clients for your real embedder, vector DB, and LLM.

Indexing

import { embedder, vectorDB } from "./clients";

interface Chunk {
  id: string;
  text: string;
  metadata: Record<string, unknown>;
}

// Recursive-ish splitter: paragraphs, packed to a max length.
function chunkDocument(doc: { id: string; text: string; title: string }): Chunk[] {
  const paras = doc.text.split(/\n\n+/).map(p => p.trim()).filter(Boolean);
  const chunks: Chunk[] = [];
  let buf = "";
  let n = 0;
  const flush = () => {
    if (!buf) return;
    chunks.push({
      id: doc.id + "-" + n++,
      text: buf,
      metadata: { source: doc.id, title: doc.title },
    });
    buf = "";
  };
  for (const p of paras) {
    if ((buf + "\n\n" + p).length > 800) flush();
    buf = buf ? buf + "\n\n" + p : p;
  }
  flush();
  return chunks;
}

async function indexDocuments(docs: { id: string; text: string; title: string }[]) {
  for (const doc of docs) {
    const chunks = chunkDocument(doc);
    const vectors = await embedder.embedBatch(chunks.map(c => c.text));
    await vectorDB.upsert(
      chunks.map((c, i) => ({ ...c, vector: vectors[i] }))
    );
  }
}

Querying

import { embedder, vectorDB, reranker, llm } from "./clients";

async function ask(question: string, userRoles: string[]): Promise<{
  answer: string;
  sources: string[];
}> {
  // 1. Embed the question (same model as indexing).
  const queryVector = await embedder.embed(question);

  // 2. Retrieve a wide candidate set, with access control at the DB.
  const candidates = await vectorDB.search(queryVector, {
    topK: 30,
    filter: { accessRoles: { in: userRoles } },
  });

  // 3. Detect "nothing relevant".
  if (candidates.length === 0) {
    return { answer: "I don't have information on that.", sources: [] };
  }

  // 4. Rerank down to the best few.
  const scores = await reranker.score(question, candidates.map(c => c.text));
  const top = candidates
    .map((c, i) => ({ ...c, score: scores[i] }))
    .sort((a, b) => b.score - a.score)
    .slice(0, 5);

  // 5. Build a grounded, citable prompt.
  const context = top
    .map((c, i) => "[" + (i + 1) + "] (" + c.metadata.source + ") " + c.text)
    .join("\n\n");

  const prompt = [
    "Answer using ONLY the context. If it's not there, say you don't know.",
    "Cite sources by number, e.g. [1].",
    "",
    "Context:",
    context,
    "",
    "Question: " + question,
  ].join("\n");

  // 6. Generate.
  const answer = await llm.complete(prompt);
  return { answer, sources: top.map(c => String(c.metadata.source)) };
}

Where to grow it

This skeleton already does chunking, embedding, access-controlled retrieval, reranking, grounding, no-answer handling, and citations. From here, layer in the upgrades from these docs in order of payoff: hybrid search, query rewriting for chat, contextual retrieval at index time, and an evaluation harness so every change is measured. Each is a self-contained addition to one of the steps above.

Tip
Build exactly this first, against your real data, and evaluate it. A tuned basic pipeline is a strong baseline — add advanced techniques only where the numbers show a gap.
Bipin Singh
Written by Bipin Singh

Senior Full-Stack Engineer · AI & AWS. I build production RAG and LLM systems for enterprises.

Work with me