iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Short answer: In Anant Kumar’s 100-question benchmark, adding an agent by itself raised exact-match accuracy only modestly, from 67% to 70%. The larger improvement came when the agent could query structured graph data: accuracy reached 99%. A typed-selection version kept that score while cutting reported latency from 13.2 seconds to 1.7 seconds per question and using no generation-model calls. That is evidence for matching the system to the task—not a universal proof that agents pay for themselves.
What the 100-question benchmark tested
Kumar reports evaluating six pipelines on 100 questions over 2,951 Wikipedia articles. The questions covered lookup, temporal, multi-hop, superlative, and aggregation tasks. Each pipeline used the same generation model, Gemini 3.1 Flash-Lite, along with local BGE embeddings and TigerGraph’s native vector index. The work was built for the TigerGraph Agentic GraphRAG Hackathon. Answers were scored by exact match against gold answers, without a model in the scoring loop. These are the author’s reported results, not independently reproduced findings. Kumar’s benchmark article
| Pipeline | Exact match | Reported tokens per question |
|---|---|---|
| RAG | 67% | 3,586 |
| GraphRAG with entity linking and one-hop traversal | 67% | 3,952 |
| Agent with text and entity tools | 70% | 6,065 |
| Agent with structured graph tools | 99% | 3,412 |
| Typed-selection planner with structured graph tools | 99% | 2,267 |
Token counts and scores are those reported by Kumar for this benchmark; they do not establish total operating cost. The text/entity agent’s three-point lift over RAG was small beside the 29-point difference between that agent and the structured-graph agent. The comparison therefore points less to “an agent is better” than to whether the system has the right representation and tools for the question. Kumar’s benchmark article
Free tools Windows power users keep installed
One-click scans. No signup required.
Why structured data changed the result
Aggregation questions expose a weakness in systems that retrieve only a limited number of passages: a top-five result can be relevant yet still omit records needed for a complete count. Kumar gives the example, “how many cycling events had more than 30 competitors?” In his setup, RAG answered 1 of 21 aggregation questions correctly, and basic GraphRAG answered 0 of 21. After parsing structured fields from Wikipedia infoboxes into an Olympic-event graph—with links to Games, Sport, and Venue, plus an edge to the previous Games—the reported result rose to 21 of 21. Kumar’s aggregation results
#1 Best Overall
That result makes sense for a count over a defined set: the system needs complete, filterable records, not merely plausible passages. It demonstrates the value of the particular graph schema on this dataset. It does not show that graph databases always beat retrieval, or that every task benefits from an agent.
When the agent’s planning may be unnecessary
The full agent with structured graph tools achieved 99% exact match at a reported average latency of 13.2 seconds per question. Kumar then replaced its generative planner with two typed selection calls. That version retained 99% exact match, reported 1.7 seconds per question, and made zero generation-model calls. Kumar’s latency and planner comparison
Rank #2
The practical distinction is the decision space. If a question can be answered by selecting among known fields, filters, or query types, a typed interface may do the job without asking a model to devise a plan. An agent is more defensible when it must adaptively investigate, decide what evidence to seek next, or combine tools in ways that are not known in advance. Kumar’s quote captures his interpretation—“agents earn their cost on open-ended questions, and they’re overkill where a typed query settles it”—but the benchmark supports it only as a hypothesis from this implementation, not a general rule.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
The article also says a 500-calls-per-day free-tier limit interrupted benchmark work and influenced interest in a no-generation route. That is a constraint in the author’s testing context, not a cost estimate for other deployments. Kumar’s account of the selection planner
What “cost” this benchmark can—and cannot—answer
The reported comparisons include accuracy, latency, token use, and generation-model calls. They do not provide complete per-question dollar costs, token prices, infrastructure expenses, or a break-even calculation. A faster system with fewer calls may reduce some expenses, but these figures alone cannot show whether it is cheaper overall or when its benefits exceed its costs.
For a real evaluation, compare approaches on representative questions and record the measures that fit your deployment:
Rank #4
- Answer accuracy against a reliable answer key, including the exact-match or other acceptance criteria that matter to users.
- Evidence completeness, especially for counts, totals, and queries that require every matching record.
- Latency, token consumption, and model-call frequency under the same workload.
- Monetary and infrastructure cost, including the services and query systems used.
- Failure detection and recovery: whether the system notices unsupported answers, missing records, or incorrect intermediate selections.
Why a high accuracy score still needs scrutiny
Kumar describes several failures that illustrate how a benchmark can mislead if its evaluation and pipeline are not checked carefully. An LLM judge rated 14 wrong answers 4 or 5 out of 5, often because they were fluent refusals. He therefore emphasized exact match and added an evidence-support verification pass. He also reports that a field-selection change caused the agent to recount a truncated evidence list, lowering exact match from 99% to 82%. A parsing bug mishandled a temporal question, and a stale benchmark artifact contained five incorrect counts; he says regression tests were added for these failures. These are author-reported details, not independently reproduced. Kumar’s account of evaluation and pipeline failures
For readers comparing systems, the lesson is to inspect not just the headline score but also the answer key, evidence path, and behavior when evidence is incomplete. A fluent answer is not proof of correctness, and a pipeline change can affect results in ways that a single final score hides.
Best Value
How to interpret the result
This benchmark is a useful case study in system design: on these questions and this corpus, agentic planning over text and entity tools offered little improvement, while structured graph access made a large difference. A typed planner then matched the structured agent’s accuracy with lower reported latency and no generation-model calls. The figures do not establish a universal threshold for agent return on investment or show that the same outcome will hold with another dataset, model, schema, or implementation.
The reported stack included TigerGraph Savanna, GSQL, and a native vector index. That identifies the author’s setup; it is not evidence that the same product or architecture is necessary for other workloads. Kumar’s implementation description
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →

