GraphRAG
Standard RAG retrieves isolated passages. It struggles with questions that require connecting facts scattered across many documents, or with "big picture" questions about a whole corpus. GraphRAG addresses this by building a knowledge graph of entities and relationships and retrieving over that structure.
The core idea
During indexing, use an LLM to extract entities (people, products, concepts) and the relationships between them from your documents, assembling a graph. At query time you can traverse relationships — following connections between facts — rather than only matching isolated text.
When it helps
- Multi-hop questions: "Which projects were led by people who previously worked at X?" — answering needs several linked facts, not one passage.
- Global / holistic questions: "What are the main themes across these 500 reports?" — no single chunk holds the answer; you need to reason over the whole set.
- Highly interconnected domains: org charts, dependencies, citations, supply chains.
Local vs global search
- Local search starts from specific entities and explores their neighbourhood in the graph — good for focused questions about particular things.
- Global search uses summaries of clusters/communities in the graph to answer broad questions about overall themes — something vanilla RAG essentially can't do.
The trade-off
GraphRAG is powerful but expensive to build: extracting entities and relationships across a corpus means many LLM calls at index time, and maintaining the graph as content changes adds complexity. Use it when your questions are genuinely about connections and themes; for straightforward "find the passage that answers this," standard RAG is simpler and cheaper.