Routing
Not every question should be answered the same way. Routing inspects the incoming query and decides where it should go — which data source, which index, which prompt, or even whether to retrieve at all. It's how a single assistant serves many kinds of question well.
Why route?
A real assistant sits in front of heterogeneous knowledge: product docs, HR policies, a SQL database of orders, live APIs. A question about last month's revenue belongs in the database, not the document store; a policy question belongs in the wiki. Blindly vector-searching everything for every query wastes effort and dilutes relevance.
Logical routing
Use an LLM (or a classifier) to categorise the query and pick a destination — a specific index, tool, or data source. Effectively a routing function whose output is a choice, not prose.
const route = await llm.complete(
"Classify the query into one of: [product_docs, hr_policy, orders_db, general].\n" +
"Reply with only the label.\n\nQuery: " + query
);
switch (route.trim()) {
case "orders_db": return answerFromSql(query);
case "hr_policy": return ragSearch(query, hrIndex);
case "product_docs": return ragSearch(query, docsIndex);
default: return generalAnswer(query);
}
Semantic routing
Instead of asking an LLM, embed the query and compare it to embeddings that represent each route (or a set of example queries per route). Pick the nearest route. Faster and cheaper than an LLM call, and often accurate enough for clear-cut categories.
Routing to "no retrieval"
Some inputs need no documents at all — greetings, chit-chat, or questions the base model can answer directly. A route that skips retrieval saves latency and cost, and avoids stuffing irrelevant context into small talk.
Prompt and model routing
Routing isn't only about data sources. Route simple queries to a small fast model and hard ones to a stronger model; route different query types to different prompt templates. This is how you keep both cost and quality under control across a wide range of inputs.