In the noisy world of AI agent hype, it is easy to assume that adding autonomous planning layers to retrieval systems yields massive performance gains. But a new benchmark conducted for the TigerGraph Agentic GraphRAG Hackathon suggests the opposite. Developer Anant Kumar measured six different retrieval pipelines against 100 evaluation questions and found that while exact match accuracy jumped from 67% to 99%, the agent architecture itself was responsible for only 3% of that improvement. The remaining 32-point leap came entirely from how the underlying data was structured, not from the agent’s reasoning capabilities.

The Illusion of Agentic Superiority

The experiment used 2,951 Wikipedia articles stored in TigerGraph Savanna, evaluated across five question types including lookup, temporal, and aggregation queries. Every pipeline utilized the same cheap generation model, gemini-3.1-flash-lite, to ensure a fair comparison. The results were stark: standard RAG and basic GraphRAG both scored 67%. Adding an agent that could plan steps and choose between five tools, including vector search and entity linking, only raised the score to 70%. It wasn’t until the agent was given access to a specifically structured graph layer that accuracy skyrocketed to 99%. The core issue with the initial setup was structural, not cognitive. Aggregation questions, such as how many cycling events had more than 30 competitors, failed because retrieval systems typically fetch the top five documents. If the answer requires counting across 43 documents, the model confidently reports the size of the retrieved set rather than the actual total. Kumar’s LLM-extracted entity graph, containing 12,000 entities with free-text labels, couldn’t solve this because the facts weren’t modeled as queryable properties. The agent’s reasoning was powerless against a flawed data foundation.

Parsing Beats Planning

The breakthrough came when Kumar stopped relying on LLMs to extract entities and instead built a deterministic parser for the Wikipedia infoboxes. This created a second graph layer with structured edges like PREV_GAMES, allowing the system to treat counting as a graph traversal rather than a retrieval problem. This ingestion method made zero LLM calls. With this structured surface, the agent’s accuracy on aggregation questions went from 0/21 to 21/21. The agent’s primary value shifted from complex reasoning to intelligent routing—switching to vector search for open-ended queries and structured queries for factual ones.

The Cost of Overkill

Perhaps the most provocative finding was that for many questions, the agent’s planning was redundant. Kumar tested replacing the generative planner with two typed selection calls, effectively removing the reasoning loop for structured data. The result was the same 99% exact match accuracy, but at 1.7 seconds per question instead of 13.2 seconds, and with zero generation-model calls. This eliminated rate limit failures and reduced token usage from 3,412 per question in the full agentic pipeline to 2,267 in the selection-based approach. The agent earns its cost only on open-ended questions; for structured facts, it’s often expensive overhead.

Key Takeaways

  • Data structure determines ceiling: Properly modeled graph edges improved accuracy by 32 points, while agent architecture alone added only 3 points.
  • Deterministic parsing outperforms LLM extraction: Building the graph from infoboxes without LLM calls solved aggregation failures that reasoning could not.
  • Agents are routing mechanisms, not just reasoners: The primary benefit was switching between structured queries and vector search, not deep logical deduction.
  • Over-engineering hurts latency: A non-generative selection planner achieved identical accuracy to the agentic planner but was nearly 8x faster and cheaper.

The Bottom Line

Stop treating agentic planning as a magic bullet for retrieval accuracy. If your data isn't structured for queryable facts, your agent is just confidently hallucinating over a broken foundation. Structure first, agentic routing second.