Build a Graph Only Where the Answer Is a Relationship
Graphs beat vector search when the answer depends on a relationship, and cost more elsewhere. Here is how to tell which case you are in.
The graph pitch this season is additive. Use GraphRAG for retrieval, a knowledge graph for reasoning, a context graph for memory, and stack all three. It sounds like maturity. Each layer looks like a capability you are missing.
If you own the roadmap for an AI feature, the question that matters is narrower: which of your failures does a graph fix, and which does it only make more expensive? A Google Cloud team published that comparison this month, on a real code-migration job. Graph RAG and standard vector RAG scored an identical 91% on CodeBLEU, the text-overlap metric most teams would have used to judge them. The structural numbers underneath were nowhere near identical.
Key takeaways
- A graph improves AI retrieval when the correct answer depends on a relationship between items, and it adds cost without benefit when the needed information sits inside one self-contained item.
- In a Google Cloud code-migration test, graph RAG cut API hallucination from 56.4% to 16.2%, while both methods scored about 91% on CodeBLEU, so text-overlap metrics hid the gap.
- Graph structure works best alongside text retrieval: in the KGCaRe study, a graph-only method scored 43.03 F1 on GPT-4o against 83.96 for vanilla RAG and 88.67 for the combined approach.
- The edge test checks whether a graph will pay: sort recent retrieval failures by whether the right answer required crossing a relationship. Only those failures justify building a graph.
- A record of why an AI decision was made is a governance deliverable with a named owner, and it should be specified before anyone chooses a graph database to store it.

Where does a graph actually beat vector search?
A graph beats vector search when the answer lives in the connection between two pieces of information rather than inside either one. The Google Cloud study, by Jaiswal and colleagues, shows this cleanly. They migrated a Java repository to Python, once with standard RAG and once with a graph built from the code’s syntax tree, with edges for inheritance, calls and imports.
On structure, the graph won decisively. Dependency and API resolution rose from 34.8% to 65.9%. Parent-child consistency rose from 26.7% to 45.5%. The true API hallucination rate fell from 56.4% to 16.2%. The mechanism is plain: semantic retrieval often fails to fetch an inherited parent class unless it happens to share wording with the child, so the model guesses.
The same paper has a table the graph enthusiasts skip. On simple, self-contained files, the entity classes that map database columns and do not heavily invoke external dependencies, the two methods scored almost the same. When the answer sits inside one file, there is no edge to follow, and a graph has nothing to add. The authors’ proposed next step is a hybrid: graph retrieval for core object-oriented logic, standard vector retrieval for unstructured documentation.

Why can’t your current evaluation see the difference?
Most teams judge retrieval with metrics that reward text that looks right, and those metrics are blind to broken relationships. In the Google Cloud study, CodeBLEU put the two approaches at 91.1% and 90.6%. The authors call it “a false equivalence”. Standard RAG produced readable Python that looked correct at a glance and failed on execution.
The same blindness hides the graph’s costs. Handing the model the whole architecture made it over-engineer: cyclomatic complexity consistency fell from 71.6% to 46.7%. The Procedural Graphs paper from Google reports a cost of its own kind: its guidance increases token use even when it reduces the number of solver steps. A graph changes what the model gets wrong. If your evaluation cannot see structure, you will not see that trade in either direction.
I’d treat this as the first decision, before any graph work. Add at least one metric that fails when a relationship is wrong: a dependency that does not resolve, a condition that was ignored, a parent the answer should have inherited from. Without it, you cannot tell whether a graph earned its keep. It is the same discipline as treating the spec as the real bottleneck: decide what correct means before you change the system.
The study’s own summary table puts the near-identical CodeBLEU scores right above the structural gaps they hide.

Exhibit 1. Aggregate metrics for standard RAG and graph RAG on the same Java-to-Python migration, from Jaiswal et al. (2026). RAG is retrieval-augmented generation; standard RAG retrieves by vector similarity, graph RAG by a code graph, and Delta is graph minus standard. Compare the first row, Avg CodeBLEU (90.6% vs 91.1%), with Avg Hallucination Rate (56.4% vs 16.2%) and Avg Dependency/API Res (DRQ) (34.8% vs 65.9%); PCC is parent-child consistency (26.7% vs 45.5%) and THC is type hint completion. Graph RAG loses on Avg Cyclomatic Consistency (CCC), 71.6% vs 46.7%, and on DP, docstring preservation, 67.0% vs 61.0%. Source: Jaiswal, N., Shukla, A., Malhotra, D., Agrawal, A., Garg, S., Bhaumik, S., & Puri, S. (2026). Beyond Vector Similarity: Hierarchical Context-Aware Graph RAG vs Standard RAG in Enterprise Code Migration. Google Cloud. arXiv:2609.12464, Table 1, p. 8.
Graph plus text beats graph alone
The strongest results come from combining graph structure with text retrieval. KGCaRe, a study from the University of Galway with Fidelity Investments, tested conditional questions that general-purpose RAG handles badly. It combined triples pulled from an automatically built knowledge graph with ordinary retrieved passages.
On 106 yes-or-no questions with GPT-4o, KGCaRe scored 88.67 F1 against 83.96 for vanilla RAG. Think-on-Graph, a graph-reasoning baseline, scored 43.03. The margin over vanilla RAG also varied by model: 7.55 points with Mistral, but one point with GPT-3.5, 71.75 against 70.75. A graph is a supplement whose value depends on the question and the model, and it should be budgeted like one.

The edge test
Before you fund a graph, run what I call the edge test. Pull your last 50 retrieval failures and sort each one by a single question: did the right answer require crossing a relationship between two items? Inheritance, dependency, a condition attached to a rule, a precedent tied to an earlier case all count. A missing fact inside one document does not.
The sort gives you a decision:
- Most failures need an edge. A graph is likely to pay. Build it for that relationship type first, and extend it only when the next failure type shows up.
- Most failures sit inside one item. Fix chunking, ranking and the evaluation metric. A graph will add cost and new failure modes without touching the problem.
- Failures are rare. An independently compiled playbook based on Andrew Ng’s agentic-design material sets a useful floor: if the failure rate is below 5%, the next pattern’s complexity cost likely exceeds its benefit.
The same playbook gives the rule I’d use to decide whether a graph, once built, is earning its place: a knowledge graph is justified when the same entity or relationship is queried by more than one agent or across more than one session. “A graph that is written to once and never queried is a database table with extra overhead.” That lines up with the harness argument in why the harness produces the capability you think you are buying: the structure around the model matters when something reads from it.
Match the graph to the question it answers
Once the edge test says yes, pick the graph by the question it has to answer. There are three, and they need different things.
What is true? That is the knowledge graph, with facts as entity-relation-entity triples. The playbook’s checklist is the part most teams skip: every edge should trace to a source document, overwrites should become supersession links rather than deletions, and entity-resolution decisions should stay inspectable.
What should the agent do next? Google’s Procedural Graph answers this with procedure-relation-procedure triples, kept outside the model weights where they can be inspected and edited without retraining. Across tasks and model families it delivered consistent gains over memory-based baselines. For many agent teams this is the graph they actually need, and it looks more like a library of reusable skills than a model of the business.
Why was this decided? This is where the “context graph” pitch overreaches. BCG’s architecture for agentic banking specifies the record in detail: for any material decision, the bank should be able to show which agent acted, under whose delegation, with which credential, on which data, through which controls and human gates, and with what result. The second line sets that evidence standard. The requirement comes first, and it has an owner. Whether it lives in a graph is an implementation detail.
So the order I’d defend in a planning review is: name the owner of the decision record, add a metric that catches broken relationships, run the edge test, and only then choose the graph. If nobody can name the owner, a third graph layer will store reasoning that nobody is accountable for.

References
- Jaiswal, N., Shukla, A., Malhotra, D., Agrawal, A., Garg, S., Bhaumik, S., & Puri, S. (2026). Beyond Vector Similarity: Hierarchical Context-Aware Graph RAG vs Standard RAG in Enterprise Code Migration. Google Cloud. arXiv:2609.12464.
- Verma, G., Sarkar, S., Pillai, D., Shiokawa, H., Xu, Y., Veazey, F., Hubbert, P., Su, H., & Buitelaar, P. (2026). KGCaRe: Explainable Complex Conditional Question Answering using Automatic Knowledge Graph Construction and Context Retrieval with LLMs. arXiv:2608.09779.
- Lu, Y., Chen, Y., Wu, S., & Arık, S. Ö. (2026). Procedural Graphs: Self-Evolving Execution Structures for LLM Agents. Google. arXiv:2609.09153.
- Graph Engineering for Multi-Agentic Systems: The Andrew Ng Playbook (2026, July). Independently compiled working note based on Andrew Ng’s courses and DeepLearning.AI material; not affiliated with or endorsed by Andrew Ng or DeepLearning.AI.
- Coppola, M., & Kleppe, A. (2026, September). When Agents Run the Bank: The End of the Second Line as You Know It. Boston Consulting Group.
Frequently asked questions
When is GraphRAG better than vector RAG?
GraphRAG is better when the correct answer depends on a relationship between items, such as inheritance, dependencies or conditions. In a Google Cloud code-migration test it cut API hallucination from 56.4% to 16.2%, while on self-contained files the two methods performed almost the same.
Why do standard metrics miss the benefit of graph RAG?
Text-overlap metrics reward output that looks right and do not penalize broken relationships. In the Google Cloud study, graph RAG and standard RAG scored 91.1% and 90.6% on CodeBLEU even though standard RAG's code failed on execution.
Should a knowledge graph replace text retrieval?
No. In the KGCaRe study, combining knowledge-graph triples with retrieved text scored 88.67 F1 on GPT-4o, while vanilla RAG scored 83.96 and a graph-reasoning baseline scored 43.03.
What is the edge test for deciding on a graph?
The edge test sorts your recent retrieval failures by whether the right answer required crossing a relationship between two items. If most failures need an edge, a graph is likely to pay; if they sit inside single items, improve chunking, ranking and evaluation instead.
Evidence
The mechanism is plain: semantic retrieval often fails to fetch an inherited parent class unless it happens to share wording with the child, so the model guesses.
When an LLM attempts to translate a child class, semantic retrieval often fails to fetch the inherited parent class or distant utility interfaces unless they share high lexical overlap.
Jaiswal, N., et al. (2026). Beyond Vector Similarity: Hierarchical Context-Aware Graph RAG vs Standard RAG in Enterprise Code Migration. Google Cloud. arXiv:2609.12464. (p. 2)
Standard RAG produced readable Python that looked correct at a glance and failed on execution.
Standard RAG’s “plain vanilla translation” behavior yields highly readable Python code that looks correct at a glance, but fails upon execution.
Jaiswal, N., et al. (2026). Beyond Vector Similarity: Hierarchical Context-Aware Graph RAG vs Standard RAG in Enterprise Code Migration. Google Cloud. arXiv:2609.12464. (p. 9)
The true API hallucination rate fell from 56.4% to 16.2%.
slashed the true API hallucination rate from a catastrophic 56.4% down to 16.2%
Jaiswal, N., et al. (2026). Beyond Vector Similarity: Hierarchical Context-Aware Graph RAG vs Standard RAG in Enterprise Code Migration. Google Cloud. arXiv:2609.12464. (p. 11)
Parent-child consistency rose from 26.7% to 45.5%.
18.8% im- provement in Parent-Child Consistency (PCC) (scoring 45.5% versus the baseline’s 26.7%)
Jaiswal, N., et al. (2026). Beyond Vector Similarity: Hierarchical Context-Aware Graph RAG vs Standard RAG in Enterprise Code Migration. Google Cloud. arXiv:2609.12464. (p. 9)
An independently compiled playbook based on Andrew Ng's agentic-design material sets a useful floor: if the failure rate is below 5%, the next pattern's complexity cost likely exceeds its benefit.
Rule 2: Measure before promoting. Be‐ fore adding the next pattern, establish a baseline with the The July 2026 debate crystallized around a real question: current pattern and measure the failure rate you expect the when does a loop stop being sufficient and a graph become new pattern to address. If the failure rate is below 5%, the necessary? The answer maps onto three criteria. First, session new pattern's complexity cost likely exceeds its benefit.
Graph Engineering for Multi-Agentic Systems: The Andrew Ng Playbook (2026). Independently compiled working note. (p. 8)
The playbook's checklist is the part most teams skip: every edge should trace to a source document, overwrites should become supersession links rather than deletions, and entity-resolution decisions should stay inspectable.
Graph Provenance Does every edge trace to a source docu‐ Arch. ment? review, and graph architecture for persistence. The com‐ Versioning Are overwrites replaced by superses‐ pound effect is where the real performance lives — not in sion links? any single pattern, but in the specific combination tuned to Entity resolu‐ Are resolution decisions inspectable?
Graph Engineering for Multi-Agentic Systems: The Andrew Ng Playbook (2026). Independently compiled working note. (p. 9)
Across tasks and model families it delivered consistent gains over memory-based baselines.
Across tasks and model families, it delivers consistent gains over memory-based baselines.
Lu, Y., Chen, Y., Wu, S., & Arık, S. Ö. (2026). Procedural Graphs: Self-Evolving Execution Structures for LLM Agents. arXiv:2609.09153. (p. 11)
Dependency and API resolution rose from 34.8% to 65.9%.
Dependency & API Resolu- tion Quality (DRQ) improved by 31.1% (from 34.8% to 65.9%)
Jaiswal, N., et al. (2026). Beyond Vector Similarity: Hierarchical Context-Aware Graph RAG vs Standard RAG in Enterprise Code Migration. Google Cloud. arXiv:2609.12464. (p. 9)
Think-on-Graph, a graph-reasoning baseline, scored 43.03.
GPT 4o Text Prompt 85.85 Think-on-Graph 43.03
Verma, G., et al. (2026). KGCaRe. arXiv:2608.09779. (p. 12)
A Google Cloud team published that comparison this month, on a real code-migration job.
Google Cloud arXiv:2609.12464v1 [cs.AI] 11 Sep 2026
Jaiswal, N., et al. (2026). Beyond Vector Similarity: Hierarchical Context-Aware Graph RAG vs Standard RAG in Enterprise Code Migration. Google Cloud. arXiv:2609.12464. (p. 1)
On 106 yes-or-no questions with GPT-4o, KGCaRe scored 88.67 F1 against 83.96 for vanilla RAG.
Table 2 Comparative analysis of KGCaRe and existing approaches on 106 Yes/No QA pairs from the ConditionalQA dataset. F1 Score LLM / Model used Approach [Yes/No type QA ] Vanilla LLM 58.49 Code Prompt 40.58 Mistral Text Prompt 30.89 Think-on-Graph 39.94 Vanilla RAG 62.26 HybridContextQA 57.13 KGCaRe (ours) 69.81 Vanilla LLM 64.22 Code Prompt 43.86 Mixtral Text Prompt 57.34 Think-on-Graph 45.65 Vanilla RAG 74.84 HybridContextQA 76.27 KGCaRe (ours) 82.07 Vanilla LLM 66.03 Code Prompt 72.07 GPT 3.5 Text Prompt 70.95 Think-on-Graph 44.34 Vanilla RAG 70.75 HybridContextQA 68.86 KGCaRe (ours) 71.75 Vanilla LLM 71.69 Code Prompt 84.90 GPT 4o Text Prompt 85.85 Think-on-Graph 43.03 Vanilla RAG 83.96 HybridContextQA 80.97 KGCaRe (ours) 88.67
Verma, G., et al. (2026). KGCaRe. arXiv:2608.09779. (p. 12)
The Procedural Graphs paper from Google reports a cost of its own kind: its guidance increases token use even when it reduces the number of solver steps.
Guidance increases token use even when it reduces solver steps
Lu, Y., Chen, Y., Wu, S., & Arık, S. Ö. (2026). Procedural Graphs: Self-Evolving Execution Structures for LLM Agents. arXiv:2609.09153. (p. 11)
The same playbook gives the rule I'd use to decide whether a graph, once built, is earning its place: a knowledge graph is justified when the same entity or relationship is queried by more than one agent or across more than one session.
Rule 5: The graph earns itself. A know‐ only record). Graphs handle all three structurally. But loops ledge graph is justified when the same entity or relationship are simpler to build, cheaper to run, and sufficient for the is queried by more than one agent or across more than one vast majority of single-agent, single-session tasks. The session.
Graph Engineering for Multi-Agentic Systems: The Andrew Ng Playbook (2026). Independently compiled working note. (p. 8)
In a Google Cloud code-migration test, graph RAG cut API hallucination from 56.4% to 16.2%, while both methods scored about 91% on CodeBLEU, so text-overlap metrics hid the gap.
dropping the API hallucination rate from 56.4% to 16.2%. Furthermore, it improves Dependency Resolution Quality (DRQ) from 34.8% to 65.9% and enhances Parent-Child Consistency (PCC) from 26.7% to 45.5%. Interest- ingly, traditional lexical metrics fail to capture this divergence; both methodologies achieved an identical 91% average CodeBLEU score
Jaiswal, N., et al. (2026). Beyond Vector Similarity: Hierarchical Context-Aware Graph RAG vs Standard RAG in Enterprise Code Migration. Google Cloud. arXiv:2609.12464. (p. 1)
On simple, self-contained files, the entity classes that map database columns and do not heavily invoke external dependencies, the two methods scored almost the same.
Low Variance Context: Files exhibiting near-zero variance between the two paradigms, such as PetType.java and Vet.java, are highly self-contained domain objects. They are simple Java POJOs (Entities) that map database columns and do not heavily invoke external business dependencies.
Jaiswal, N., et al. (2026). Beyond Vector Similarity: Hierarchical Context-Aware Graph RAG vs Standard RAG in Enterprise Code Migration. Google Cloud. arXiv:2609.12464. (p. 10)
The second line sets that evidence standard.
Evidence & Reporting Continuous, provable Generates lineage Sets evidence L7 (Explainability, Traceability, record of how decisions & evidence standards; reporting
Coppola, M., & Kleppe, A. (2026). When Agents Run the Bank. Boston Consulting Group. (p. 5)
The authors' proposed next step is a hybrid: graph retrieval for core object-oriented logic, standard vector retrieval for unstructured documentation.
Future research will focus on mitigating these trade-offs by developing dynamic Hybrid RAG architectures. By intelligently orchestrating between deterministic Graph RAG for core object-oriented logic and standard Vector RAG for unstructured documentation
Jaiswal, N., et al. (2026). Beyond Vector Similarity: Hierarchical Context-Aware Graph RAG vs Standard RAG in Enterprise Code Migration. Google Cloud. arXiv:2609.12464. (p. 11)
Graph structure works best alongside text retrieval: in the KGCaRe study, a graph-only method scored 43.03 F1 on GPT-4o against 83.96 for vanilla RAG and 88.67 for the combined approach.
Vanilla LLM 71.69 Code Prompt 84.90 GPT 4o Text Prompt 85.85 Think-on-Graph 43.03 Vanilla RAG 83.96 HybridContextQA 80.97 KGCaRe (ours) 88.67
Verma, G., et al. (2026). KGCaRe. arXiv:2608.09779. (p. 12)
The margin over vanilla RAG also varied by model: 7.55 points with Mistral, but one point with GPT-3.5, 71.75 against 70.75.
Vanilla LLM 58.49 Code Prompt 40.58 Mistral Text Prompt 30.89 Think-on-Graph 39.94 Vanilla RAG 62.26 HybridContextQA 57.13 KGCaRe (ours) 69.81 Vanilla LLM 64.22 Code Prompt 43.86 Mixtral Text Prompt 57.34 Think-on-Graph 45.65 Vanilla RAG 74.84 HybridContextQA 76.27 KGCaRe (ours) 82.07 Vanilla LLM 66.03 Code Prompt 72.07 GPT 3.5 Text Prompt 70.95 Think-on-Graph 44.34 Vanilla RAG 70.75 HybridContextQA 68.86 KGCaRe (ours) 71.75
Verma, G., et al. (2026). KGCaRe. arXiv:2608.09779. (p. 12)
In the Google Cloud study, CodeBLEU put the two approaches at 91.1% and 90.6%.
metrics like CodeBLEU—which scored a near-identical 91.1% and 90.6% for both paradigms
Jaiswal, N., et al. (2026). Beyond Vector Similarity: Hierarchical Context-Aware Graph RAG vs Standard RAG in Enterprise Code Migration. Google Cloud. arXiv:2609.12464. (p. 11)
Google's Procedural Graph answers this with procedure-relation-procedure triples, kept outside the model weights where they can be inspected and edited without retraining.
The PG keeps a task domain’s procedural knowledge outside the model weights, where it can be inspected, retrieved at each step, and edited without retraining.
Lu, Y., Chen, Y., Wu, S., & Arık, S. Ö. (2026). Procedural Graphs: Self-Evolving Execution Structures for LLM Agents. arXiv:2609.09153. (p. 2)
They migrated a Java repository to Python, once with standard RAG and once with a graph built from the code's syntax tree, with edges for inheritance, calls and imports.
Utilizing tree-sitter, we deterministically extract polyglot Abstract Syntax Trees (AST) and map their architectural dependencies (e.g., INHERITS, CALLS, IMPORTS) into a Google Cloud Spanner Property Graph.
Jaiswal, N., et al. (2026). Beyond Vector Similarity: Hierarchical Context-Aware Graph RAG vs Standard RAG in Enterprise Code Migration. Google Cloud. arXiv:2609.12464. (p. 2)
Graph RAG and standard vector RAG scored an identical 91% on CodeBLEU, the text-overlap metric most teams would have used to judge them.
Interest- ingly, traditional lexical metrics fail to capture this divergence; both methodologies achieved an identical 91% average CodeBLEU score
Jaiswal, N., et al. (2026). Beyond Vector Similarity: Hierarchical Context-Aware Graph RAG vs Standard RAG in Enterprise Code Migration. Google Cloud. arXiv:2609.12464. (p. 1)
Handing the model the whole architecture made it over-engineer: cyclomatic complexity consistency fell from 71.6% to 46.7%.
Providing the LLM with dense, global structural context introduces new vulnerabili- ties: Graph RAG suffers a severe degradation in Cyclomatic Complexity Consistency (dropping from Standard RAG’s 71.6% to 46.7%) due to defensive over-engineering by the LLM
Jaiswal, N., et al. (2026). Beyond Vector Similarity: Hierarchical Context-Aware Graph RAG vs Standard RAG in Enterprise Code Migration. Google Cloud. arXiv:2609.12464. (p. 1)
"A graph that is written to once and never queried is a database table with extra overhead."
A graph that is written to once and never queried is a engineering decision is not "which is better" but "when does database table with extra overhead.
Graph Engineering for Multi-Agentic Systems: The Andrew Ng Playbook (2026). Independently compiled working note. (p. 8)
BCG's architecture for agentic banking specifies the record in detail: for any material decision, the bank should be able to show which agent acted, under whose delegation, with which credential, on which data, through which controls and human gates, and with what result.
For any material agentic decision, the bank should have the ability to show which agent acted, under whose delegation, The engine should use deterministic logic wherever a case fits a using which credential, on which data, through which known schema, reserving flexible reasoning for exceptions when governed services, Policy-as-Code controls, and human the system cannot proceed safely. The second line sets the judgment gates, and with what result.
Coppola, M., & Kleppe, A. (2026). When Agents Run the Bank. Boston Consulting Group. (p. 6)
The authors call it "a false equivalence".
This creates a false equivalence, demonstrating that CodeBLEU exhibits a strong bias toward localized string overlap and is incapable of penalizing a module for severely broken API contracts or hallucinated imports.
Jaiswal, N., et al. (2026). Beyond Vector Similarity: Hierarchical Context-Aware Graph RAG vs Standard RAG in Enterprise Code Migration. Google Cloud. arXiv:2609.12464. (p. 9)
KGCaRe, a study from the University of Galway with Fidelity Investments, tested conditional questions that general-purpose RAG handles badly.
Answering complex conditional questions using Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) remains a challenge, particularly in domain-specific contexts where general-purpose LLMs and RAG tend to underperform.
Verma, G., et al. (2026). KGCaRe. arXiv:2608.09779. (p. 1)
It combined triples pulled from an automatically built knowledge graph with ordinary retrieved passages.
The relevant triples extracted from the KG in path form, along with semantically retrieved text passages, are then fed into custom KGCaRe prompts to generate answers to the complex conditional questions
Verma, G., et al. (2026). KGCaRe. arXiv:2608.09779. (p. 2)