Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Docling can miss a table, merge table cells, scramble page columns, or assign every PDF heading level 1 for different reasons. Start by identifying which stage failed: OCR, layout detection, table structure, reading order, or heading hierarchy. Then change only the relevant setting and compare the result with the source page.

First identify what Docling got wrong

Docling processes documents through distinct stages. Layout detection identifies regions such as text, tables, pictures, and section headers; OCR supplies text when a usable PDF text layer is absent; table-structure recognition reconstructs cells within detected table regions; later stages handle reading order and heading hierarchy. A setting for one stage will not necessarily fix another.

  • No table region detected: investigate layout detection and the page content, not just table reconstruction.
  • A table is detected but its cells are wrong: test table structure mode and cell matching.
  • Prose flows across page columns: investigate reading order, not table settings.
  • Headings are recognized but all have the same depth: enable heading hierarchy inference.
  • Text is missing or garbled: check whether the PDF has a reliable text layer and whether OCR is needed.

Docling’s model catalog describes separate layout, OCR, and table-recognition components, including TableFormer fast and accurate modes and OCR options such as Tesseract, EasyOCR, RapidOCR, macOS Vision, and SuryaOCR. Availability depends on the installed release and environment.

Why Docling misses a table

Check whether the page has usable text

Determine whether the PDF is digitally generated, scanned, or mixed. Compare extracted text with the visible page. If text is missing or garbled, check OCR availability and language configuration for the affected pages. OCR is not automatically beneficial for a digital PDF with a sound text layer. Inspect OCR results against the original, especially for small type, symbols, and dense numeric tables.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
Epson Workforce ES-50 Compact & Lightweight Mobile Document Scanner
  • PORTABLE SCANNER FOR USE ON-THE-GO — The fastest and lightest mobile single-sheet-fed compact document scanner in its class¹
  • QUICK DOCUMENT SCANNING ― This Epson ultra-fast scanner scans a single page as quickly as 5.5 seconds²; Windows and Mac compatible
  • VERSATILE PAPER HANDLING ― Portable scanner scans documents up to 8.5 x 72 in; Also easily digitizes receipts and ID cards to make accounting, bookkeeping, and organizing simpler
  • INTUITIVE, HIGH-SPEED SOFTWARE — Epson ScanSmart Software³ is a smart tool allowing you to easily scan, review, and save; Stay organized easily with the help of this Epson scanner
  • EASY SETUP — USB-powered connect to your computer for quick and simple scanning; No batteries or external power supply required to operate portable document scanner; Standard Connectivity: USB 2.0

Check whether layout detected a table region

Table structure recognition works on regions supplied by layout analysis. If the visible table was not identified as a table, changing the table reconstruction mode cannot reliably recover it. Inspect the page representation and determine whether the region was labeled as a table before adjusting structure settings.

For a detected but poorly reconstructed table, use accurate mode

Docling’s advanced options documentation describes TableFormer accurate mode as the default and intended for difficult table structures; fast mode trades some accuracy for speed. Defaults may differ in older or customized pipelines, so check the configuration for your installed release. Accurate mode can help with structure reconstruction, but it does not guarantee a correct result or fix a table region that layout analysis missed.

For merged columns inside a table, test cell matching

By default, Docling maps recognized table structure back to PDF cells. The documentation says disabling cell matching can improve cases where multiple table columns merge. Compare both outputs on the affected page:

Rank #2
Sale
Brother DS-640 Compact Mobile Document Scanner, (Model: DS640)
  • FAST SPEEDS - Scans color and black and white documents a blazing speed up to 16ppm (1). Color scanning won’t slow you down as the color scan speed is the same as the black and white scan speed.
  • ULTRA COMPACT – At less than 1 foot in length and only about 1. 5lbs in weight you can fit this device virtually anywhere (a bag, a purse, even a pocket).
  • READY WHENEVER YOU ARE – The DS-640 mobile scanner is powered via an included micro USB 3. 0 cable allowing you to use it even where there is no outlet available. Plug it into you PC or laptop and you are ready to scan.
  • WORKS YOUR WAY – Use the Brother free iPrint&Scan desktop app for scanning to multiple “Scan-to” destinations like PC, Network, cloud services, Email and OCR. (2) Supports Windows, Mac and Linux and TWAIN/WIA for PC/ICA for Mac/SANE drivers. (3)
  • OPTIMIZE IMAGES AND TEXT – Automatic color detection/adjustment, image rotation (PC only), bleed through prevention/background removal, text enhancement, color drop to enhance scans. Software suite includes document management and OCR software. (4)
pipeline_options.table_structure_options.do_cell_matching = False

This setting concerns cells inside an extracted table. It is not a general fix for text flowing across ordinary page columns.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Check intentionally blank columns separately

A Docling project discussion records a maintainer explanation that post-processing removes fully empty rows and columns because they may be prediction anomalies; a June 2025 reply reported that behavior still occurring. Do not assume accurate mode preserves an intentionally empty column. Compare the exported table with the source in the exact release you use.

Why prose columns are read in the wrong order

Newspaper-style or multi-column page text is a reading-order problem, not necessarily a table problem. Enabling table structure or disabling cell matching does not ensure that prose is read down one column before moving to the next.

Rank #3
Sale
Epson Workforce ES-400 II High-Speed Color Duplex Desktop Document Scanner
  • FAST DOCUMENT SCANNING — Document scanner with feeder allows you to speed through stacks with a 50-sheet Auto Document Feeder (ADF); Efficient office scanner to help you scan more productively
  • INTUITIVE, HIGH-SPEED SOFTWARE — Quickly scan with this desktop document scanner; Epson ScanSmart Software lets you easily preview scans, email files, upload to the cloud, and more; Plus, automatic file naming saves even more time
  • SEAMLESS INTEGRATION — Easily incorporate your data into most document management software with the included TWAIN driver; Office document scanner integrates seamlessly with business workflows
  • EASY SHARING — Duplex scanner allows you to scan straight to email or popular cloud storage2 services like Dropbox, Evernote, Google Drive, and OneDrive for simple storage and sharing
  • SIMPLE FILE MANAGEMENT — Scanner allows the creation of searchable PDFs with Optical Character Recognition (OCR) and convert scans to editable Word or Excel files effortlessly; Designed for home and office document scanning

Docling’s rule-based reading-order stage can use visible PDF rules as signals and is enabled by default. If rules separating columns or horizontal bands appear to disrupt ordering, test disabling those separators:

pipeline_options = PdfPipelineOptions(use_reading_order_separators=False)

The CLI equivalent is --no-reading-order-separators. The documentation says this option changes ordering only. A Docling issue filed for version 2.43.0 describes text flowing across columns in a three-column financial document, despite table-related settings. It is an example of a reported failure, not evidence that all current versions behave the same way.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why PDF headings are all level 1

Detecting a section header and determining its place in a hierarchy are separate tasks. Docling’s layout model can mark a block as a section header without deciding whether it is, for example, a top-level heading or a subsection. The advanced options documentation explains that this is why PDF headings are level 1 by default.

Rank #4
Canon Canoscan Lide 300 Scanner (PDF, AUTOSCAN, Copy, Send)
  • Scanner type: Document
  • Connectivity technology: USB
  • With Auto Scan Mode, the scanner automatically detects what you're scanning
  • Digitize documents and images

Enable heading hierarchy inference in the PDF pipeline:

from docling.datamodel.pipeline_options import HeadingHierarchyOptions, PdfPipelineOptions

pipeline_options = PdfPipelineOptions()
pipeline_options.heading_hierarchy_options = HeadingHierarchyOptions(enabled=True)
pipeline_options.generate_parsed_pages = True

The hierarchy stage infers levels using PDF bookmarks first, then heading numbering, then visual style such as size, weight, slant, and case. It changes section-header levels; it does not add or reorder document content. Parsed pages are needed for font-style inference.

Scanned pages have a limitation: OCR does not provide font weight or slant metadata, so visual-style inference ranks headings by size alone. Bookmarks or numbering can still provide separate signals when present. Review inferred levels against the page instead of assuming a scan preserves all the styling cues of a digital PDF. See the heading-level documentation for the configuration and inference details.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Sale
ScanSnap iX2500 Wireless or USB High-Speed Document Scanner, Black
  • OUR MOST ADVANCED SCANSNAP. Large touchscreen, fast 45ppm double-sided scanning, 100-sheet document feeder, Wi-Fi and USB connectivity, automatic optimizations, and support for cloud services. Upgraded replacement for the discontinued iX1600
  • CUSTOMIZABLE. SHARABLE. Select personalized profiles from the touchscreen. Send to PC, Mac, mobile devices, and clouds. QUICK MENU lets you quickly scan-drag-drop to your favorite computer apps
  • STABLE WIRELESS OR USB CONNECTION. Built-in Wi-Fi 6 for the fastest and most secure scanning. Connect to smart devices or cloud services without a computer. USB-C connection also available
  • PHOTO AND DOCUMENT ORGANIZATION MADE EFFORTLESS. Easily manage, edit, and use scanned data from documents, receipts, photos, and business cards. Automatically optimize, name, and sort files
  • AVOIDS PAPER JAMS AND DAMAGE. Features a brake roller system to feed paper smoothly, a multi-feed sensor that detects pages stuck together, and skew detection to prevent paper damage and data loss
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Minimal PDF configurations to compare

These examples reflect the documented API, but names and defaults can change. Check the documentation for your installed Docling release before using them in a production pipeline.

Difficult table structure

from docling.datamodel.base_models import InputFormat
from docling.datamodel.pipeline_options import PdfPipelineOptions, TableFormerMode
from docling.document_converter import DocumentConverter, PdfFormatOption

pipeline_options = PdfPipelineOptions(do_table_structure=True)
pipeline_options.table_structure_options.mode = TableFormerMode.ACCURATE
converter = DocumentConverter(
    format_options={InputFormat.PDF: PdfFormatOption(pipeline_options=pipeline_options)}
)

Heading hierarchy

from docling.datamodel.pipeline_options import HeadingHierarchyOptions, PdfPipelineOptions

pipeline_options = PdfPipelineOptions()
pipeline_options.heading_hierarchy_options = HeadingHierarchyOptions(enabled=True)
pipeline_options.generate_parsed_pages = True

Visible rules disrupting reading order

pipeline_options = PdfPipelineOptions(use_reading_order_separators=False)

Validate the fix against the page

Review the same affected pages after each targeted change. Compare the visible page with extracted text, table cells, reading order, and heading levels as separate outputs. For important documents, keep representative pages as a regression sample so you can check whether a configuration change or upgrade alters results.

Docling documents parsed-page and image controls for inspecting the underlying page representation. If you consider another OCR or document-parsing approach for pages that remain unreliable, compare it on the same representative pages: check support for the document type and OCR languages, treatment of page columns versus tables, handling of merged cells and blank columns, traceability to page regions, and whether processing is local or remote. Docling’s documentation says remote OCR and hosted-model use is explicitly opt-in; decide whether that data flow is appropriate for the documents before enabling a remote service.

Quick Recap

Bestseller No. 4
Canon Canoscan Lide 300 Scanner (PDF, AUTOSCAN, Copy, Send)
Canon Canoscan Lide 300 Scanner (PDF, AUTOSCAN, Copy, Send)
Scanner type: Document; Connectivity technology: USB; With Auto Scan Mode, the scanner automatically detects what you're scanning
$75.00

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.