RAG Handbook retrieval-augmented generation Bipin Singh
The core pipeline

Loading & ingestion

2 min readChapter 04 of 26By Bipin Singh

Ingestion is the unglamorous first step where most real-world RAG projects actually stall. Before anything can be chunked or embedded, you have to get clean text out of messy sources — and the quality of that text sets a hard ceiling on everything downstream.

Where content comes from

Parsing is the hard part

Raw formats hide traps. A PDF is a layout format, not a text format — text can arrive in the wrong reading order, tables collapse into gibberish, and multi-column pages interleave. HTML is full of nav bars, cookie banners, and ads that pollute your corpus. Getting this right matters more than any clever retrieval trick later.

Watch out
Tables and figures are where naive pipelines quietly lose information. A table flattened into a run-on sentence retrieves poorly and reads worse. Consider extracting tables separately (as Markdown or HTML) and, for image-heavy content, using a vision model to describe figures.

Cleaning

Normalise the extracted text so downstream steps behave consistently:

Metadata is not optional

For every chunk, capture where it came from. Metadata powers filtering, access control, citations, and debugging. At minimum store:

{
  text: "The refund window is 30 days from delivery...",
  metadata: {
    source: "policies/returns.pdf",
    title: "Returns & Refunds Policy",
    section: "Refund eligibility",
    page: 3,
    url: "https://intranet/policies/returns",
    createdAt: "2025-02-01",
    tenantId: "acme",          // for multi-tenant isolation
    accessRoles: ["support"]    // for access control at query time
  }
}
Tip
Decide your metadata schema before you index at scale. Backfilling metadata means re-indexing everything. Source, title, section, timestamp, and any tenant/permission keys are the ones you'll wish you had.

Plan for change from day one

Content is not static. You need a story for adding, updating, and deleting documents without rebuilding the whole index. Give every source a stable ID and a content hash so you can detect what changed and re-embed only that. Deletions matter too — for correctness and for the right to be forgotten. See Production & scaling for incremental indexing patterns.

Bipin Singh
Written by Bipin Singh

Senior Full-Stack Engineer · AI & AWS. I build production RAG and LLM systems for enterprises.

Work with me