How RAG works
RAG has an offline half and an online half. The offline half prepares your knowledge; the online half answers questions with it. Here is the whole flow.
Phase 1 — Indexing (offline)
This runs ahead of time, in a batch or on a schedule, whenever your source content changes.
- Load. Pull raw content from its source: files, a database, a website, an API.
- Parse & clean. Extract usable text from PDFs, HTML, Word docs, and so on; strip boilerplate.
- Chunk. Split long documents into passages small enough to retrieve precisely but large enough to carry meaning.
- Embed. Run each chunk through an embedding model to get a vector — a list of numbers that captures its meaning.
- Store. Save the vectors, the original text, and any metadata in a vector database.
Phase 2 — Retrieval & generation (online)
This runs per request, in the hot path, and must be fast.
- Embed the query. Convert the user's question into a vector using the same embedding model you used for the chunks.
- Search. Find the chunks whose vectors are closest to the query vector — the nearest neighbours in meaning.
- Augment. Assemble a prompt that contains the question plus the retrieved chunks as context.
- Generate. Send the prompt to the LLM, which writes an answer grounded in the supplied context.
- Return. Deliver the answer, ideally with citations back to the source chunks.
A minimal end-to-end sketch
Stripped of frameworks, the online path looks like this:
async function answer(question, db, llm, embedder) {
// 1. Embed the question with the same model used for indexing.
const queryVector = await embedder.embed(question);
// 2. Retrieve the most similar chunks.
const hits = await db.search(queryVector, { topK: 5 });
// 3. Build a grounded prompt.
const context = hits.map((h, i) => "[" + (i + 1) + "] " + h.text).join("\n\n");
const prompt =
"Answer the question using ONLY the context below. " +
"If the answer is not in the context, say you do not know.\n\n" +
"Context:\n" + context + "\n\n" +
"Question: " + question;
// 4. Generate.
return llm.complete(prompt);
}
The data flow at a glance
Reading left to right: content becomes vectors once; questions become vectors every time and are matched against them.
INDEXING (offline, batch)
docs -> parse -> chunk -> embed -> [vector DB]
QUERY (online, per request)
question -> embed --------> [vector DB] --search--> top-k chunks
|
question + chunks --> LLM --> answer (+ citations)
Where quality is won or lost
Newcomers assume the LLM is the hard part. In practice, retrieval quality dominates. If the right passage never makes it into the prompt, no model — however capable — can answer correctly. That is why most of these docs are about the parts around the LLM: chunking, embeddings, search, and reranking. Garbage retrieved, garbage generated.