RAG Handbook retrieval-augmented generation Bipin Singh
Advanced RAG

Query transformations

2 min readChapter 11 of 26By Bipin Singh

Users ask messy, short, ambiguous questions. The raw query is often a poor search key. Query transformation improves retrieval by rewriting or expanding the question before it hits the vector database — frequently the highest-return upgrade to naive RAG.

Why transform the query?

The user's phrasing may not match how the answer is written. A follow-up like "and the cost?" carries no standalone meaning. A broad question may need several targeted searches. Transformations bridge the gap between how people ask and how documents are written.

Query rewriting

Use an LLM to turn a raw or conversational query into a clean, standalone search query — resolving pronouns, expanding abbreviations, and folding in context from the chat history. Essential for multi-turn assistants.

// Turn a context-dependent follow-up into a standalone query.
const standalone = await llm.complete(
  "Given the conversation, rewrite the final user message as a " +
  "self-contained search query.\n\n" +
  "Conversation:\n" + history + "\n\nFollow-up: " + userMessage
);
const hits = await db.search(await embedder.embed(standalone));

Multi-query

Generate several paraphrases of the question, search with each, and merge the results (deduplicated, often via RRF). This casts a wider net so a single unlucky phrasing doesn't sink retrieval. Costs a few extra embeddings and searches; noticeably improves recall.

HyDE — hypothetical document embeddings

Counter-intuitive but effective: ask the LLM to write a hypothetical answer to the question, then embed that answer and search with it. The reasoning is that an answer looks more like the target document than the question does — so its embedding lands closer to real answer passages. The hypothetical may contain errors; that's fine, you only use its embedding to retrieve real, correct passages.

Tip
HyDE shines when questions and source documents are written very differently — e.g. terse user questions against formal, verbose documentation.

Step-back prompting

For narrow or detailed questions, first ask a more general "step-back" question to retrieve broader background, then combine that context with results from the specific query. Helps when the answer requires foundational context the specific query alone wouldn't surface.

Decomposition

Complex, multi-part questions ("compare our refund policy with our warranty terms") retrieve badly as one blob. Break them into sub-questions, retrieve and answer each, then synthesise. This turns one hard retrieval into several easy ones.

"Compare refund and warranty policies for defective items"
   -> "What is the refund policy for defective items?"
   -> "What is the warranty policy for defective items?"
   -> retrieve + answer each, then synthesise a comparison

The cost trade-off

Every transformation adds LLM calls and latency before retrieval even starts. Apply them where they pay off — multi-turn chat almost always needs rewriting; simple one-shot lookups may need nothing. Adaptive systems classify the query first and only transform when it's worth it.

Bipin Singh
Written by Bipin Singh

Senior Full-Stack Engineer · AI & AWS. I build production RAG and LLM systems for enterprises.

Work with me