RAG Handbook retrieval-augmented generation Bipin Singh
Advanced RAG

Routing

2 min readChapter 13 of 26By Bipin Singh

Not every question should be answered the same way. Routing inspects the incoming query and decides where it should go — which data source, which index, which prompt, or even whether to retrieve at all. It's how a single assistant serves many kinds of question well.

Why route?

A real assistant sits in front of heterogeneous knowledge: product docs, HR policies, a SQL database of orders, live APIs. A question about last month's revenue belongs in the database, not the document store; a policy question belongs in the wiki. Blindly vector-searching everything for every query wastes effort and dilutes relevance.

Logical routing

Use an LLM (or a classifier) to categorise the query and pick a destination — a specific index, tool, or data source. Effectively a routing function whose output is a choice, not prose.

const route = await llm.complete(
  "Classify the query into one of: [product_docs, hr_policy, orders_db, general].\n" +
  "Reply with only the label.\n\nQuery: " + query
);

switch (route.trim()) {
  case "orders_db":    return answerFromSql(query);
  case "hr_policy":    return ragSearch(query, hrIndex);
  case "product_docs": return ragSearch(query, docsIndex);
  default:             return generalAnswer(query);
}

Semantic routing

Instead of asking an LLM, embed the query and compare it to embeddings that represent each route (or a set of example queries per route). Pick the nearest route. Faster and cheaper than an LLM call, and often accurate enough for clear-cut categories.

Routing to "no retrieval"

Some inputs need no documents at all — greetings, chit-chat, or questions the base model can answer directly. A route that skips retrieval saves latency and cost, and avoids stuffing irrelevant context into small talk.

Prompt and model routing

Routing isn't only about data sources. Route simple queries to a small fast model and hard ones to a stronger model; route different query types to different prompt templates. This is how you keep both cost and quality under control across a wide range of inputs.

Tip
Start without routing. Add it once you clearly have distinct sources or query types that need different handling — premature routing is complexity you'll pay to maintain before it earns its keep.
Bipin Singh
Written by Bipin Singh

Senior Full-Stack Engineer · AI & AWS. I build production RAG and LLM systems for enterprises.

Work with me