Building & reference
Reference implementation
Here is a complete, deliberately minimal RAG pipeline in TypeScript — no framework, so every step is visible. It's a teaching skeleton, not production code, but it maps one-to-one to the concepts in these docs. Swap the placeholder clients for your real embedder, vector DB, and LLM.
Indexing
import { embedder, vectorDB } from "./clients";
interface Chunk {
id: string;
text: string;
metadata: Record<string, unknown>;
}
// Recursive-ish splitter: paragraphs, packed to a max length.
function chunkDocument(doc: { id: string; text: string; title: string }): Chunk[] {
const paras = doc.text.split(/\n\n+/).map(p => p.trim()).filter(Boolean);
const chunks: Chunk[] = [];
let buf = "";
let n = 0;
const flush = () => {
if (!buf) return;
chunks.push({
id: doc.id + "-" + n++,
text: buf,
metadata: { source: doc.id, title: doc.title },
});
buf = "";
};
for (const p of paras) {
if ((buf + "\n\n" + p).length > 800) flush();
buf = buf ? buf + "\n\n" + p : p;
}
flush();
return chunks;
}
async function indexDocuments(docs: { id: string; text: string; title: string }[]) {
for (const doc of docs) {
const chunks = chunkDocument(doc);
const vectors = await embedder.embedBatch(chunks.map(c => c.text));
await vectorDB.upsert(
chunks.map((c, i) => ({ ...c, vector: vectors[i] }))
);
}
}
Querying
import { embedder, vectorDB, reranker, llm } from "./clients";
async function ask(question: string, userRoles: string[]): Promise<{
answer: string;
sources: string[];
}> {
// 1. Embed the question (same model as indexing).
const queryVector = await embedder.embed(question);
// 2. Retrieve a wide candidate set, with access control at the DB.
const candidates = await vectorDB.search(queryVector, {
topK: 30,
filter: { accessRoles: { in: userRoles } },
});
// 3. Detect "nothing relevant".
if (candidates.length === 0) {
return { answer: "I don't have information on that.", sources: [] };
}
// 4. Rerank down to the best few.
const scores = await reranker.score(question, candidates.map(c => c.text));
const top = candidates
.map((c, i) => ({ ...c, score: scores[i] }))
.sort((a, b) => b.score - a.score)
.slice(0, 5);
// 5. Build a grounded, citable prompt.
const context = top
.map((c, i) => "[" + (i + 1) + "] (" + c.metadata.source + ") " + c.text)
.join("\n\n");
const prompt = [
"Answer using ONLY the context. If it's not there, say you don't know.",
"Cite sources by number, e.g. [1].",
"",
"Context:",
context,
"",
"Question: " + question,
].join("\n");
// 6. Generate.
const answer = await llm.complete(prompt);
return { answer, sources: top.map(c => String(c.metadata.source)) };
}
Where to grow it
This skeleton already does chunking, embedding, access-controlled retrieval, reranking, grounding, no-answer handling, and citations. From here, layer in the upgrades from these docs in order of payoff: hybrid search, query rewriting for chat, contextual retrieval at index time, and an evaluation harness so every change is measured. Each is a self-contained addition to one of the steps above.
Tip
Build exactly this first, against your real data, and evaluate it. A tuned basic pipeline is a strong baseline — add advanced techniques only where the numbers show a gap.