Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Docling can extract tables from a PDF into pandas DataFrames, which you can save as CSV files for Excel. Its official Python example demonstrates CSV and HTML export—not creation of an .xlsx workbook. To produce a real Excel workbook, add a separate pandas or Excel-writing step after extraction.

What Docling does—and what it does not export

Docling converts a PDF into a document representation whose tables can be iterated and exported as pandas DataFrames. The official table-export example writes those DataFrames to CSV and demonstrates HTML export. CSV opens in Excel, but it is not an Excel workbook: it does not preserve workbook features such as multiple sheets or cell formatting.

The extraction flow is therefore distinct from the final file-format choice: convert the PDF and extract tables with Docling, then save the resulting data as CSV or use an additional workbook-writing step for .xlsx. The cited example does not document the latter step.

Extract tables to CSV with Docling

The following pattern follows the documented API: convert the PDF, iterate over the converted document’s tables, export each to a DataFrame, and save each DataFrame as a separate CSV file.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from pathlib import Path
from docling.document_converter import DocumentConverter

result = DocumentConverter().convert("input.pdf")
output_dir = Path("tables")
output_dir.mkdir(exist_ok=True)

for i, table in enumerate(result.document.tables, start=1):
    df = table.export_to_dataframe(doc=result.document)
    df.to_csv(output_dir / f"table-{i}.csv", index=False)

This code expects Docling and pandas to be installed; both are named as prerequisites in the official example. Installation commands and package versions are not specified here, so use the current Docling installation guidance rather than relying on a version-specific command copied from an older tutorial.

Each extracted table gets a file such as table-1.csv. The index=False argument avoids writing the DataFrame row index as an extra CSV column. If a rendered table view is useful, the same example also shows HTML export.

Save extracted data as a real Excel workbook

If your requirement is specifically an .xlsx file, CSV is only a handoff format. After producing each DataFrame, write it to a workbook using a separate Excel-writing library or a pandas Excel writer. That workbook-writing step is outside the cited Docling example; choose and configure it according to your environment and output needs.

For a workbook with one sheet per extracted table, the basic shape is to open one Excel writer and send each DataFrame to a sheet. Sheet names must meet Excel’s naming rules, so short names such as Table1 are safer than deriving sheet names directly from arbitrary PDF content. The example below shows the separate pandas step; it is not a Docling feature.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import pandas as pd

with pd.ExcelWriter("tables.xlsx") as writer:
    for i, table in enumerate(result.document.tables, start=1):
        df = table.export_to_dataframe(doc=result.document)
        df.to_excel(writer, sheet_name=f"Table{i}", index=False)

Use the CSV loop when separate, simple files suit the workflow. Use a workbook writer when users need multiple tables organized as sheets in one .xlsx file.

Choose table-recognition settings when structure is difficult

Extracted values depend on the PDF’s layout and the structure-recognition options; there is no accuracy guarantee for every document. Docling’s table-structure documentation describes settings that affect how predicted table structure is mapped to text cells:

  • Cell matching: do_cell_matching controls whether structure predictions are mapped back to text cells found in the PDF. The documentation notes that using structure-predicted text cells can improve results when multiple columns have been merged erroneously.
  • Recognition mode: TableFormerMode.FAST is faster but less accurate, while TableFormerMode.ACCURATE is the more accurate option for difficult structures and the documented default. These are tradeoffs, not fixes guaranteed to work on a particular PDF.

When a table’s columns are misaligned or merged, compare the DataFrame with the PDF and try the documented alternatives. Check the current documentation for the configuration syntax supported by the Docling release you have installed.

Handle scanned PDFs and OCR separately

A scanned or image-only PDF requires text recognition as well as table-structure recognition. Docling’s CLI reference exposes OCR engine choices and a table-recognition switch. OCR determines what text can be recognized from page images; table recognition determines how that content is organized into cells. Neither should be treated as a substitute for checking the extracted result.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The cited sources do not establish a best OCR engine or benchmark one against another. Test the available configuration on representative pages from your own PDFs, particularly where scans are faint, skewed, or contain dense tables.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Check the output against the PDF

Before using extracted values for analysis or reporting, compare the output to the source page. Prioritize checks for:

  • Column boundaries, especially where the extraction has merged adjacent columns.
  • Scanned pages, where OCR may miss or misread characters.
  • Multi-level or hierarchical tables. A Docling community discussion reports that indentation and formatting cues may not carry through as label hierarchy in DataFrame or Markdown output. Treat that as a reason to inspect these tables, not as a universal behavior.
  • Row and column counts, totals, and representative cell values, including any blank cells or repeated headers that affect downstream work.

If a mismatch matters, correct the DataFrame or adjust the recognition configuration and regenerate the output. Do not assume that a successful conversion means every cell was interpreted correctly.

Pick the output and settings that fit your task

Decision Option What it means
File format CSV Shown in the official Docling example; opens in Excel, but is not a workbook.
File format .xlsx Requires a separate workbook-writing step; creation is not demonstrated in the cited Docling example.
Text-cell mapping do_cell_matching Controls mapping predictions to PDF text cells; the table-structure documentation describes alternatives and their potential effect.
Table structure mode FAST Faster, with lower accuracy than ACCURATE according to Docling’s table-structure documentation.
Table structure mode ACCURATE More accurate for difficult structures and the documented default, according to the same documentation.
Input type Scanned or image-only PDF Consider OCR configuration separately from table recognition; the CLI reference lists OCR engine choices.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.