Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

RAG systems often get table questions wrong because retrieving a few similar text chunks is not the same as preserving a table or calculating across all its rows. GraphRAG can add useful entity relationships and corpus-wide context, but it is not a table calculator. For exact sums, counts, filters, percentages, and comparisons, keep the data structured and use SQL; use retrieval for the surrounding prose.

Why does RAG get table values wrong?

A table’s meaning depends on relationships among its headers, rows, cells, units, and sometimes footnotes. When a system flattens a table into text and splits it into chunks, those relationships can be disrupted. A retrieved chunk might contain a number without the header that explains it, or only a portion of the relevant rows.

This creates several distinct failure points:

  • Retrieval failure: the needed rows or context are not retrieved.
  • Representation failure: flattening or chunking separates values from headers or other structural context.
  • Execution failure: the system does not reliably calculate over the complete set of relevant rows.
  • Generation failure: the model states more than the retrieved evidence supports.

These are relevant risks, not a universal explanation for every incorrect answer. The 2025 TableRAG paper describes structural information loss and a lack of a global view in heterogeneous-document question answering. Its example includes a percentage calculated over retrieved top-N chunks rather than the full table. It does not establish a general rate of table-related RAG hallucinations.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should you use SQL or GraphRAG for a table question?

Choose the mechanism based on the operation the question requires. If it asks for an exact result across rows, make the calculation a structured operation rather than asking a language model to infer it from partial passages.

Approach Best fit Strength Important limit
Baseline vector RAG A question answerable from a few relevant passages Simple top-k text retrieval; GraphRAG also provides basic search May miss aggregation or retrieve fragmented table context. [TableRAG paper; Microsoft GraphRAG overview]
GraphRAG local search Entity-specific questions involving connected concepts and source text Combines graph-derived context with related source text Not documented as exact SQL calculation over arbitrary tables. [Microsoft query documentation]
GraphRAG global search Corpus-wide themes or synthesis Uses community reports and map-reduce synthesis Resource-intensive; summaries do not replace exact table execution. [Microsoft global search documentation]
Structured table store + SQL and text retrieval Exact filters, counts, sums, percentages, or cross-row calculations mixed with document context SQL operates on structured data; the TableRAG paper describes a hybrid text-and-SQL design Requires table extraction, schema handling, and query validation. [TableRAG paper]

A practical design is to keep tables in a database, retrieve relevant prose separately, execute the tabular subtask with validated SQL, and compose the results with their source context. This follows the text-plus-SQL approach described by TableRAG; it is not a performance guarantee for every implementation.

For a simple lookup where headers are clear and only a small amount of context is needed, carefully preserved table markup or row serialization may be enough. For calculations over all relevant rows, do not rely on a top-k text sample as though it were the full table.

What does GraphRAG add?

Microsoft’s GraphRAG workflow extracts entities, relationships, and claims from text units, clusters the entity graph into communities, and generates community summaries. At query time, graph structure and summaries augment the prompt. The project overview cautions that using GraphRAG out of the box may not give the best results and recommends prompt tuning.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Local search

Use local search when the question centers on a specific entity and its connected entities, relationships, and associated source text. It combines graph-derived context with text from source documents.

Global search

Use global search for broad themes or patterns across a corpus. It synthesizes community reports in a map-reduce fashion, and Microsoft documents this mode as resource-intensive.

DRIFT and basic search

DRIFT search starts from an entity and uses community context to broaden and refine exploration. Basic search retains ordinary top-k vector retrieval for questions that baseline RAG can answer adequately.

These modes organize and retrieve context; the documentation does not describe them as guaranteed exact arithmetic over arbitrary tables. Keep exact tabular operations in a structured execution path.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to start GraphRAG locally

Microsoft’s documented quickstart is a local CLI and project workflow, not an offline-only model setup. The documentation specifies Python 3.10–3.12. Its documented OpenAI or Azure OpenAI configuration requires an API key for model calls, and GraphRAG can consume substantial LLM resources. Start with a small, representative corpus.

  1. Create and enter a project directory:
    mkdir graphrag_quickstart
    cd graphrag_quickstart
  2. Create and activate a virtual environment, then install GraphRAG:
    python -m venv .venv
    source .venv/bin/activate          # Unix/macOS
    python -m pip install graphrag
  3. Initialize the workspace:
    graphrag init

    Initialization creates project files including settings.yaml, an input directory, and a generated .env file. Set the API key there for the documented OpenAI or Azure OpenAI route. Review the model and pipeline settings, then place a small text corpus in input/.

  4. Build the index:
    graphrag index

    The default index produces Parquet outputs and embeddings in the configured vector store.

  5. Try a broad query and an entity-focused query:
    graphrag query "What are the top themes in this corpus?"
    graphrag query "Which entities are connected to the key subject?" --method local

    The default query example uses global search; the second command explicitly selects local search.

    Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For table-heavy questions, add a separate data path: load table sources into a structured database, preserve links from records to their source document and table, route sums, counts, filters, and cross-row comparisons through validated SQL, and retrieve explanatory text separately. Compose the answer from the query result and source context. This is an architectural recommendation informed by TableRAG’s hybrid approach, not a recipe tested by the GraphRAG quickstart.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the TableRAG benchmark does—and does not—show

The 2025 TableRAG paper introduces HeteQA, a benchmark with 304 examples across nine domains and five tabular operations per example. Those figures describe the benchmark’s scope; they are not a general hallucination rate or evidence that GraphRAG fixes table errors. The paper concerns heterogeneous-document question answering, so its findings should not be generalized beyond that experimental setting.

For a production system, evaluate with representative questions from your own documents. Include lookups, filters, aggregations, percentages, and comparisons; verify that the right rows and headers reach the calculation step, and check that answers distinguish computed results from explanatory prose.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.