Modern AI Engineering

Lesson 10.12 · 27 min

GraphRAG: Combining Knowledge Graphs with Retrieval

Ask a normal RAG bot "What are the main themes across all 5,000 of our incident reports?" and it will summarise five random chunks. How could it ever see the whole picture?

In short: GraphRAG uses an LLM to read every chunk of a corpus and extract entities and relationships into a knowledge graph, groups the graph into communities of closely connected entities, and writes a summary for each community. At question time, local search starts from the entities a question mentions and gathers their neighbourhood, while global search combines community summaries to answer questions about the whole dataset. It answers connection and big-picture questions that chunk retrieval cannot, at a much higher indexing cost.

What is GraphRAG?

GraphRAG is a family of RAG methods that build and use a knowledge graph of the content. A knowledge graph stores knowledge as nodes (entities: people, teams, products, systems, places, events) connected by edges (relationships: "leads", "depends on", "reports to"), each usually with a short text description. The name became widely known through Microsoft Research's GraphRAG work, described in the 2024 paper "From Local to Global: A Graph RAG Approach to Query-Focused Summarization" (Edge et al.) and released as open source.

Our running example: an engineering organisation's internal wiki and incident reports. People lead projects, projects use databases and queues, incidents affect services. Questions like "If Postgres goes down, which project leads should we alert?" or "What are the recurring causes of our incidents this year?" are about connections and the whole collection, not about one paragraph.

Think of it like a detective's evidence board Normal RAG is a detective searching a filing cabinet for the page that best matches a keyword. GraphRAG first pins every person, place and event on a board and draws string between related ones. Questions like "who is connected to whom?" or "what are the main clusters of activity?" become easy once the board exists; building the board is the hard work.

Why normal RAG is not enough

Normal (vector) RAG embeds chunks and retrieves the few most similar to the question. That is excellent for local, single-passage questions: "What is the on-call policy?". It struggles with two kinds of questions:

  • Connection (multi-hop) questions: "Which project leads depend on Postgres?" The fact "Falcon stores data in Postgres" and the fact "Asha leads Falcon" may sit in different documents that share no words with each other. Similarity search on the question may find one and miss the other.
  • Global (sensemaking) questions: "What are the main themes in these reports?" No single chunk contains the answer; it is spread across the entire corpus. Retrieving the top 5 chunks summarises 5 chunks, not 5,000 reports.

The GraphRAG paper frames global questions as query-focused summarisation of a whole dataset, and that is where it showed the clearest gains over standard vector RAG, especially on how comprehensive and diverse the answers were.

Pause and think: Which question is a good fit for plain vector RAG, and which needs something like GraphRAG: (a) "What is the refund window?" (b) "What themes come up most across all customer complaints?"

(a) is a local question answered by one passage; vector RAG handles it well. (b) is a global question whose answer is spread across the whole corpus; GraphRAG's community summaries and global search are designed for it.

The big picture of GraphRAG

GraphRAG has two phases. An expensive indexing phase turns raw text into a graph with summaries. A query phase uses that structure to assemble the right context for the LLM.

How GraphRAG builds the knowledge graph

Indexing, step by step

  1. Split into text units: Documents are chunked into units of a few hundred to over a thousand tokens.
  2. LLM extraction: For each unit, an LLM is prompted to list entities (name, type, description) and relationships (source, target, description, strength). Some setups also extract claims, such as "incident #42 was caused by a bad config push".
  3. Merge and summarise: The same entity found in many chunks ("Postgres", "PostgreSQL") is merged into one node; its many descriptions are summarised into one.
  4. Community detection: A graph algorithm finds groups of nodes that are more connected to each other than to the rest. Running it hierarchically gives big communities that split into smaller sub-communities.
  5. Community reports: An LLM writes a summary for each community at each level, from the nodes, edges and claims inside it.
  6. Embed for lookup: Entity descriptions (and often text units and reports) are embedded so queries can find their starting entities.

A community is a cluster of entities with many links among themselves and few links outside. In our example, Asha, Ben, Project Falcon and Postgres form one; Carla, Dev, Project Owl and Kafka another. A community report might read: "Project Falcon, led by Asha with Ben, stores its data in Postgres." These pre-written summaries are what make global questions answerable.

tiny_graphrag.py

from collections import Counter, defaultdict
# Step 1 (normally done by an LLM reading every chunk): entity-relation triples
triples = [("Asha", "leads", "Falcon"), ("Ben", "works on", "Falcon"),
("Falcon", "stores data in", "Postgres"), ("Ben", "reports to", "Asha"),
("Carla", "leads", "Owl"), ("Dev", "works on", "Owl"),
("Owl", "streams events via", "Kafka"), ("Dev", "reports to", "Carla"),
("Owl", "reads reports from", "Postgres"), ("Eve", "leads", "Hawk"),
("Hawk", "caches with", "Redis"), ("Fay", "works on", "Hawk")]
graph = defaultdict(set)                       # Step 2: build the graph
for s, _, o in triples:
graph[s].add(o); graph[o].add(s)
# Step 3: find communities (GraphRAG uses Leiden; label propagation is simpler)
label = {n: n for n in sorted(graph)}
for _ in range(5):
for n in sorted(graph):
counts = Counter(label[m] for m in graph[n])
top = max(counts.values())
label[n] = min(l for l, c in counts.items() if c == top)
communities = defaultdict(list)
for n, l in label.items():
communities[l].append(n)
for l, members in communities.items():
print("community:", sorted(members))
# Local search: start at an entity and walk 2 hops, collecting facts
def local(entity, hops=2):
seen, frontier = {entity}, {entity}
for _ in range(hops):
frontier = {m for n in frontier for m in graph[n]} - seen
seen |= frontier
return [f"{s} {r} {o}" for s, r, o in triples if s in seen and o in seen]
facts = local("Postgres")
print(f"local('Postgres'): {len(facts)} of {len(triples)} facts reached")
for fact in facts:
print("  ", fact)

Output:

community: ['Asha', 'Ben', 'Falcon', 'Postgres']
community: ['Carla', 'Dev', 'Kafka', 'Owl']
community: ['Eve', 'Fay', 'Hawk', 'Redis']
local('Postgres'): 9 of 12 facts reached
Asha leads Falcon
Ben works on Falcon
Falcon stores data in Postgres
Ben reports to Asha
Carla leads Owl
Dev works on Owl
Owl streams events via Kafka
Dev reports to Carla
Owl reads reports from Postgres

The algorithm found the three teams as communities without being told. Local search from Postgres reached 9 of 12 facts: both projects that use it (Falcon, Owl), their leads (Asha, Carla) and team members. That is exactly the context needed for "which project leads should we alert if Postgres goes down?", even though no single document states it. The Hawk team, which uses Redis, was correctly left out.

How GraphRAG answers a question: local vs global search

Local search is for questions about specific entities. The query is matched (usually by embedding similarity) to entity descriptions to find starting nodes. The system gathers their neighbours, the relationships between them, the original text units that mention them, and relevant community reports, ranks and trims all of this to fit the context window, and asks the LLM to answer.

Global search is for questions about the whole dataset. It uses a map-reduce pattern over community reports at a chosen level: in the map step, the LLM reads each report (or batch of reports) and writes a partial answer with a helpfulness score; in the reduce step, the most helpful partial answers are combined into one final answer. Because every community contributes, the answer reflects the whole corpus, not just the top few chunks.

When to use GraphRAG

  • Use it for corpora full of interlinked entities (organisations, people, systems, cases, research literature) where questions ask about relationships, dependencies or paths.
  • Use it when users ask global, summarising questions about a large collection: themes, trends, recurring causes.
  • Use it when explainability matters: an answer can point to the entities and relationships it used.
  • Skip it when questions are mostly local look-ups ("what does policy X say?"); vector or hybrid RAG is cheaper and usually just as good.
  • Skip it when data changes constantly and re-indexing cost is unaffordable, or when the corpus is small enough to read whole.

Where it is used Intelligence and investigative analysis (who is connected to whom), enterprise knowledge bases with many teams and systems, research-literature exploration, IT dependency and incident analysis, and compliance reviews over large document sets. Open-source tooling includes Microsoft's graphrag package, plus graph features in LlamaIndex, LangChain and graph databases such as Neo4j. Approaches and APIs are evolving quickly; check current docs.

Trade-offs of GraphRAG

Costs and risks (qualitative; actual numbers depend heavily on corpus size and model choice).
AspectVector RAGGraphRAG
Indexing workOne embedding per chunkSeveral LLM calls per chunk plus summaries for every community
Indexing costLowHigh; can be orders of magnitude more than embedding
UpdatesRe-embed changed chunksRe-extract, re-merge, possibly re-cluster and re-summarise
Local questionsStrongStrong, with multi-hop links
Global questionsWeakStrong (global search)
Failure modeMissed or irrelevant chunksExtraction errors, merged or duplicated entities

Common mistakes Running GraphRAG on a corpus where questions are simple look-ups (paying a lot for nothing). Not reviewing extraction quality: if the LLM merges "Apple (company)" with "apple (fruit)" or misses relationships, every later step inherits the error. Using generic entity types that do not fit the domain; tune the extraction prompt. Using global search for specific questions (slow and vague) or local search for themes (narrow).

Researchers are actively reducing the indexing cost. For example, Microsoft described LazyGraphRAG (late 2024), which defers most LLM summarisation until query time, making indexing much cheaper. Expect this area to keep changing.

Worked example, step by step

The trade-offs table says GraphRAG indexing is "high" cost. Let us count the model calls for the 5,000 incident reports from the hook. The setup below is illustrative; real counts depend on the configuration and on the data.

Counting calls

  1. Text units: 5,000 reports × 2 text units each = 10,000 units.
  2. Vector RAG index: 10,000 embedding calls, and the index is done.
  3. GraphRAG extraction: One LLM call per unit plus one extra "gleaning" pass to catch missed entities: 20,000 LLM calls.
  4. Entity summaries: Say 3,000 entities show up in more than one unit and need their descriptions merged: 3,000 more LLM calls.
  5. Community reports: Say the graph has 400 communities across all levels: 400 more LLM calls.
  6. Total: About 23,400 LLM calls against 10,000 embedding calls. And each LLM call reads and writes far more tokens than an embedding call.
  7. Query time: A global question over a level with 120 community reports, read in batches of 10: 12 map calls + 1 reduce call = 13 LLM calls. A local question needs 1.
Model calls in this illustrative setup.
StageVector RAGGraphRAG
Build the index10,000 embedding callsAbout 23,400 LLM calls
Local question1 LLM call1 LLM call
Global question1 LLM call that sees 5 chunks13 LLM calls that see every community
One report editedRe-embed 2 unitsRe-extract 2 units; refresh the entity and community summaries they touch

One failure is worth checking before any of this money is spent at full scale: entity merging. If the extractor creates "Postgres", "PostgreSQL" and "the PG cluster" as three separate nodes, the facts about one database are split three ways, and a local search from any one node misses the rest. Two cheap checks on a sample: list all node names in alphabetical order and look for near-duplicates, and list the nodes that have exactly one edge.

Practice: try it yourself

The earlier code did local search. Now we build a toy global search: a map step that turns each community into a partial answer, and a reduce step that merges the partial answers into one. We ask "what are the most common causes of incidents?" and compare the result with what a plain top-k retriever would see. Simple counting stands in for the LLM in both steps.

practice_global_search.py

from collections import Counter
# Incident causes, already grouped into communities by the graph step.
# Illustrative data: each string is the cause noted in one incident report.
communities = {
"payments": ["bad config push", "expired certificate", "bad config push",
"database failover"],
"search":   ["memory leak", "bad config push", "memory leak"],
"mobile":   ["expired certificate", "third-party outage", "bad config push"],
}
def map_step(name, causes):          # stand-in for an LLM reading one community report
points = Counter(causes)
return {"community": name, "points": points, "helpfulness": len(causes)}
def reduce_step(partials, top=3):    # stand-in for the LLM merging partial answers
total = Counter()
for part in sorted(partials, key=lambda p: -p["helpfulness"]):
total.update(part["points"])
return total.most_common(top)
partials = [map_step(name, causes) for name, causes in communities.items()]
for part in partials:
print(f"map {part['community']:8} -> {dict(part['points'])}")
print("global answer ->", reduce_step(partials))
# What plain top-k retrieval sees: only the few chunks nearest to the question.
top_k_chunks = communities["search"]            # pretend these 3 chunks matched best
print("top-3 chunks  ->", Counter(top_k_chunks).most_common(3))
print("LLM calls: global =", len(partials) + 1, "| top-k = 1")

Output:

map payments -> {'bad config push': 2, 'expired certificate': 1, 'database failover': 1}
map search   -> {'memory leak': 2, 'bad config push': 1}
map mobile   -> {'expired certificate': 1, 'third-party outage': 1, 'bad config push': 1}
global answer -> [('bad config push', 4), ('expired certificate', 2), ('memory leak', 2)]
top-3 chunks  -> [('memory leak', 2), ('bad config push', 1)]
LLM calls: global = 4 | top-k = 1

Now change it:

  • In reduce_step, keep only the two most helpful partial answers by adding [:2] after the sorted(...) call. Predict whether the top theme changes, and which community gets dropped.
  • Change top_k_chunks to communities["payments"][:3]. Predict what top-k reports now, and decide whether it was right for a good reason or by luck.
  • Add a fourth community, "billing": ["expired certificate"] * 3. Predict the new number one theme in the global answer.

Pause and think: The top-3 chunks say "memory leak" is the main cause. The global answer says "bad config push" with 4 incidents. The three chunks were all relevant to the question, so why did top-k get it wrong?

It only saw one corner of the data. "Bad config push" appears once or twice in every community, so no single community makes it look dominant; it only stands out when all partial answers are added together. Themes that are spread thinly across the whole corpus are exactly what top-k retrieval misses and what map-reduce over communities finds.

Pause and think: In our code the reduce step merges every partial answer, so sorting by helpfulness changes nothing. Why does a real system still need that score?

Because of the context limit. With hundreds of community reports, the reduce step cannot read every partial answer, so it takes the most helpful ones first and drops the rest. If the helpfulness scores are poor, useful information is thrown away before the final answer is written.

Quick summary

  • GraphRAG builds a knowledge graph of entities and relationships from text with an LLM.
  • It groups the graph into communities and writes a summary report for each.
  • Local search starts from the question's entities and gathers their neighbourhood: good for multi-hop questions.
  • Global search map-reduces community reports: good for whole-dataset questions that vector RAG cannot answer.
  • The price is a much more expensive, slower-to-update index; use it where connections and big-picture questions matter.

Pause and think: Our team asks "What are the three biggest recurring causes of outages this year?" over 4,000 incident reports. Should we use local or global search, and why?

Global search. The answer is spread across the whole collection, and no specific entity anchors the question. Map-reduce over community summaries lets every cluster of incidents contribute before the results are combined.

Key takeaways

  • GraphRAG extracts entities and relationships with an LLM to build a knowledge graph of the corpus.
  • Community detection groups related entities; LLM-written community reports summarise each group.
  • Local search expands from the question's entities through the graph, handling multi-hop questions.
  • Global search map-reduces community reports to answer whole-dataset questions vector RAG cannot.
  • Indexing is far more expensive and slower to update than vector RAG, so use GraphRAG where connections and big-picture questions matter.

Key terms

  • Knowledge graph: A network of entities (nodes) linked by typed relationships (edges).
  • Entity extraction: Using a model to find people, systems, places and other things, plus their relations, in text.
  • Community: A group of graph nodes more densely linked to each other than to the rest of the graph.
  • Community report: An LLM-written summary of the entities, relationships and key facts in one community.
  • Local search: Answering by expanding from entities mentioned in the question through their graph neighbourhood.
  • Global search: Answering by map-reducing community reports so the whole dataset contributes.
  • Leiden algorithm: A community-detection algorithm used by Microsoft GraphRAG to find hierarchical communities.

← 10.11 Agentic RAG: Dynamic Retrieval with Multi-Step Reasoning · 10.13 Vectorless RAG: Retrieval Without Embeddings or a Vector Store →