RAG Handbook retrieval-augmented generation Bipin Singh
Advanced RAG

Self-correcting & agentic RAG

2 min readChapter 14 of 26By Bipin Singh

Basic RAG is a straight line: retrieve once, generate once. Advanced architectures add loops and self-checks so the system can notice bad retrieval, fix it, and try again — trading latency and cost for reliability.

Self-RAG

The model decides whether retrieval is needed, and after generating, critiques its own output — is it supported by the evidence? is it relevant? — and can retrieve again or revise. Retrieval becomes conditional and self-assessed rather than automatic.

Corrective RAG (CRAG)

Add a lightweight grader that evaluates retrieved documents before generation. If they look relevant, proceed. If they look weak, take corrective action — reformulate the query, fall back to web search, or strip out the irrelevant parts. The point is to catch a bad retrieval before it poisons the answer.

retrieve -> grade documents
   |-- relevant   -> generate
   |-- ambiguous  -> refine query, retrieve again
   |-- irrelevant -> fall back (e.g. web search), then generate

Adaptive RAG

Route by query complexity. A simple factual lookup gets a single retrieve-and-generate pass; a complex multi-hop question gets decomposition, multiple retrievals, and synthesis; trivial questions skip retrieval entirely. You spend effort in proportion to difficulty instead of treating every query the same.

RAG-Fusion

Combine multi-query generation with rank fusion: expand the question into several variants, retrieve for each, and fuse the ranked lists (typically with RRF) into one strong result set. A simple, effective recall booster.

Agentic RAG

The most flexible pattern: an agent treats retrieval as one tool among several and reasons in a loop — plan, act, observe, repeat — until it can answer. It can search multiple sources, call APIs, run calculations, re-query based on what it learned, and reflect on whether it has enough.

Watch out
Loops and agents cost more calls, more latency, and more ways to fail. Reach for them when accuracy genuinely requires multi-step reasoning — not by default. A well-tuned linear pipeline beats a flaky agent for most FAQ-style workloads.

Choosing an architecture

If you need...Consider
Reliable single-fact answersTuned basic RAG + reranking
To catch bad retrievalsCRAG (document grading)
Effort matched to difficultyAdaptive RAG
Higher recall cheaplyRAG-Fusion / multi-query
Multi-hop, multi-source reasoningAgentic RAG
Bipin Singh
Written by Bipin Singh

Senior Full-Stack Engineer · AI & AWS. I build production RAG and LLM systems for enterprises.

Work with me