Loading & ingestion
Ingestion is the unglamorous first step where most real-world RAG projects actually stall. Before anything can be chunked or embedded, you have to get clean text out of messy sources — and the quality of that text sets a hard ceiling on everything downstream.
Where content comes from
- Files: PDF, DOCX, PPTX, HTML, Markdown, plain text, CSV.
- Structured stores: SQL and NoSQL databases, spreadsheets, data warehouses.
- SaaS & APIs: Notion, Confluence, Google Drive, Slack, Jira, GitHub, Zendesk.
- The web: crawled or scraped pages, sitemaps, RSS feeds.
Parsing is the hard part
Raw formats hide traps. A PDF is a layout format, not a text format — text can arrive in the wrong reading order, tables collapse into gibberish, and multi-column pages interleave. HTML is full of nav bars, cookie banners, and ads that pollute your corpus. Getting this right matters more than any clever retrieval trick later.
- PDFs: simple text extractors work for clean documents; scanned or complex PDFs need OCR and layout-aware parsing. Tables and figures often need special handling.
- HTML: extract the main content and drop boilerplate (headers, footers, menus). Preserve headings and lists — they carry structure you'll want when chunking.
- Office docs: keep heading hierarchy and table structure where you can.
Cleaning
Normalise the extracted text so downstream steps behave consistently:
- Collapse repeated whitespace and fix broken line breaks (PDFs love to split sentences across lines).
- Remove page numbers, running headers/footers, and watermarks.
- Normalise unicode (quotes, dashes, ligatures) so identical text embeds identically.
- Drop empty or near-empty fragments.
Metadata is not optional
For every chunk, capture where it came from. Metadata powers filtering, access control, citations, and debugging. At minimum store:
{
text: "The refund window is 30 days from delivery...",
metadata: {
source: "policies/returns.pdf",
title: "Returns & Refunds Policy",
section: "Refund eligibility",
page: 3,
url: "https://intranet/policies/returns",
createdAt: "2025-02-01",
tenantId: "acme", // for multi-tenant isolation
accessRoles: ["support"] // for access control at query time
}
}
Plan for change from day one
Content is not static. You need a story for adding, updating, and deleting documents without rebuilding the whole index. Give every source a stable ID and a content hash so you can detect what changed and re-embed only that. Deletions matter too — for correctness and for the right to be forgotten. See Production & scaling for incremental indexing patterns.