GraphRAG
GraphRAG is a family of retrieval-augmented generation approaches that derives a graph of entities and relationships from a source corpus and uses graph structure, clusters, or summaries to support model answers. Microsoft's reference pipeline extracts a knowledge graph, builds a hierarchy of communities, produces summaries, and offers query modes for local and corpus-wide questions. Other implementations may use different graph stores and retrieval strategies.
Origin and context
Microsoft Research publicly introduced GraphRAG in February 2024 for connecting information and summarizing themes across narrative private datasets. The April paper then formalized global questions that require synthesis across an entire corpus, a difficult case for baseline RAG that retrieves a few semantically similar chunks. Microsoft subsequently documented an open pipeline, while Neo4j published an independent GraphRAG package that integrates graph retrieval with several model providers.
Why it matters
Flat chunk retrieval is effective when the question maps to a small number of passages, but it can miss distributed themes and multi-hop relationships. A graph can preserve explicit connections and provide higher-level summaries, helping analysts explore who or what is connected and what patterns span a corpus. GraphRAG is especially relevant for document collections where relationships, entities, and corpus-level sensemaking matter more than isolated passage lookup.
Example
A risk team could process incident reports into entities such as suppliers, systems, locations, and failure types, then connect co-occurring or extracted relationships. Local search could answer which incidents involve one supplier; global search could summarize recurring failure patterns across communities. Analysts should retain links from graph nodes and summaries back to source passages so they can verify an answer against the underlying reports.
How it differs
Retrieval-Augmented Generation
RAG is the broad pattern of retrieving external evidence for generation. GraphRAG adds graph construction and graph-aware retrieval or summarization. It is a RAG specialization, not a replacement term for every retrieval system that stores metadata or links.
Long-context language models
Long context supplies more raw material directly to a model. GraphRAG preprocesses a corpus into relationships and summaries, then selects graph-derived evidence. The approaches can be combined, but each introduces different costs and failure modes.
Maturity and evidence
Maturity is rated 3. GraphRAG has a documented research method, an open Microsoft implementation, and an independent Neo4j implementation. However, graph extraction, community detection, query modes, and evaluation remain implementation-dependent, and production evidence is less mature than for conventional RAG.
Limits and open questions
Graph construction adds model calls, storage, latency, and update complexity before users can query the corpus. Entity resolution and relation extraction can create false or duplicate nodes; community summaries can omit minority evidence or propagate an early error. Global answers may be expensive, and benefits depend on the question type. Teams should benchmark against simpler RAG, preserve provenance, measure extraction quality, and define incremental rebuild procedures.
Related terms
References
- From Local to Global: A Graph RAG Approach to Query-Focused SummarizationarXiv · 2024-04-24 · class A
- GraphRAG documentationMicrosoft · 2024 · class A
- Neo4j GraphRAG for Python documentationNeo4j · 2024 · class A
- GraphRAG: Unlocking LLM discovery on narrative private dataMicrosoft Research · 2024-02-13 · class A
Last updated: 2026-08-27
This term is also covered in the Skills Atlas as graphrag skill.
This term is also covered in the Skills Atlas as knowledge graphs skill.
This term is also covered in the Skills Atlas as retrieval augmented generation skill.