Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Docling can convert PDFs to readable Markdown or structured JSON, with options for table reconstruction, OCR, and reading-order handling. Use Markdown when people will read the result; choose JSON when software needs to inspect document elements and their sequence. Table and reading-order results depend on the PDF’s layout, so validate representative pages against the original.

Convert a PDF with Docling

Docling’s PDF support includes layout, reading order, and table understanding. The quickest way to try it is the command-line interface:

docling convert report.pdf --to md

This writes Markdown output. To request structured JSON instead, use:

docling convert report.pdf --to json

Both formats come from Docling’s document representation; they differ in how you consume the result. See the Docling homepage for the command examples and supported formats.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Canon Canoscan Lide 300 Scanner (PDF, AUTOSCAN, Copy, Send)
  • Scanner type: Document
  • Connectivity technology: USB
  • With Auto Scan Mode, the scanner automatically detects what you're scanning
  • Digitize documents and images

Choose Markdown or JSON

Output Best suited to What it gives you
Markdown Human-readable results, documentation, or review Readable text with tables rendered in Markdown form.
JSON Programmatic processing or inspection of document structure Structured elements, including text, tables, pictures, key-value items, and a body tree.

For downstream code, JSON is useful because DoclingDocument represents more than a flat text string. Its body tree preserves hierarchy and sequence: the concept documentation says, “The reading order of the document is encapsulated through the body tree and the order of children in each item in the tree.” See DoclingDocument concepts.

Extract tables and choose a table mode

Table structure extraction is controlled by the PDF pipeline option do_table_structure. In Python, configure a converter with DocumentConverter, PdfFormatOption, and PdfPipelineOptions; the official examples show enabling table extraction and exporting Markdown:

Rank #2
Sale
Brother DS-640 Compact Mobile Document Scanner, (Model: DS640)
  • FAST SPEEDS - Scans color and black and white documents a blazing speed up to 16ppm (1). Color scanning won’t slow you down as the color scan speed is the same as the black and white scan speed.
  • ULTRA COMPACT – At less than 1 foot in length and only about 1. 5lbs in weight you can fit this device virtually anywhere (a bag, a purse, even a pocket).
  • READY WHENEVER YOU ARE – The DS-640 mobile scanner is powered via an included micro USB 3. 0 cable allowing you to use it even where there is no outlet available. Plug it into you PC or laptop and you are ready to scan.
  • WORKS YOUR WAY – Use the Brother free iPrint&Scan desktop app for scanning to multiple “Scan-to” destinations like PC, Network, cloud services, Email and OCR. (2) Supports Windows, Mac and Linux and TWAIN/WIA for PC/ICA for Mac/SANE drivers. (3)
  • OPTIMIZE IMAGES AND TEXT – Automatic color detection/adjustment, image rotation (PC only), bleed through prevention/background removal, text enhancement, color drop to enhance scans. Software suite includes document management and OCR software. (4)
from docling.document_converter import DocumentConverter, PdfFormatOption
from docling.datamodel.base_models import InputFormat
from docling.datamodel.pipeline_options import PdfPipelineOptions

pipeline_options = PdfPipelineOptions()
pipeline_options.do_table_structure = True

converter = DocumentConverter(
    format_options={
        InputFormat.PDF: PdfFormatOption(pipeline_options=pipeline_options)
    }
)
result = converter.convert("report.pdf")
markdown = result.document.export_to_markdown()
print(markdown)

Docling’s examples describe a faster approximate table mode and an accurate mode using TableFormer for complex tables, including merged cells. The choice is a trade-off between processing needs and table complexity, not a universal guarantee about the result. Check the official examples and PDF pipeline options for the current configuration details.

Validate reconstructed tables

Compare extracted output with the source page, especially when a table has:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Plustek PS186 Desktop Document Scanner, with 50-Pages Auto Document Feeder (ADF). for Windows 7/8 / 10/11 (Intel/AMD only)
  • Up to 255 customize favorite scan file setting with "Single Touch" , Support Windows 7/8/10
  • Turn paper documents into searchable, editable files - save scans as searchable PDF files; OCR function included
  • Info Barcode function - automatic categorization of complicate documentation and data with 1D or 2D Barcode page.
  • Intelligent color and image adjustments — Auto Rotate, Crop, Deskew and blank page remove with Plustek Image Processing Technology
  • Easy send scanned files to FTP server or personal NAS (FTP) with PDFs , Jpeg , TIFF or Png format. User can download scanner driver from Plustek website
  • Headers that span columns or rows
  • Merged cells or irregular row boundaries
  • Captions or explanatory text positioned close to the table

Confirm that rows and columns retain their intended relationships and that captions have not been absorbed into table content or detached from it. Treat this as a document-specific check; the documentation does not establish a general extraction-accuracy rate.

Preserve and inspect reading order

In JSON, inspect the body tree and the order of its child items to understand Docling’s represented reading sequence. The model also distinguishes page furniture, such as headers and footers, from the main body. Markdown can be useful for a quick visual review, while JSON exposes the hierarchy directly.

Rank #4
Hczrc Portable Scanner, Photo Scanner for A4 Documents, Handheld Scanner for Business, Photo, Picture, Receipts, Books, JPG/PDF Format Selection, UP to 900 DPI, with 16G SD Car
  • Note: No software installation is required. You need 2 AA batteries ( not included) and a memory card ( included) to use it directly. Scan mode: Press and hold "Scan" for 2 seconds to turn on the device, and then press "Scan", the green light is on. The scanner moves to scan the file until the green light turns off automatically (or press the "Scan" key and the green light goes out). The number shown on the display increases by 1 to indicate that the scan is complete.
  • Portable Scanner scans images or pictures quickly: Store JPEG/PDF files within seconds, scan images or pictures quickly, plug and play, no need any software preinstalled. Compatible with Windows XP/7/Vista/Mac OS 10.4 or above version.
  • Lightweight and travel-friendly: Stored in Micro SD card directly, support read data on your computer or phone with USB connected. Powered by 2pcs AA batteries, Compact Design, it is convenient to carry outside.
  • 3 Image Resolution: 3 modes of resolution for your options: 300dpi/600dpi/900dpi, you can save it at the clearest way, picture and document are showed clear as it is. Freely choose your favorite resolution.File Format: JPEG/PDF format is all available, Great storage capacity as it supports 32G Micro SD card(Included 16GB Card),total meet your need for business trip or daily use.
  • Widely Used: It is applicable in bank, insurance business, real estate agency,home, office, library or outdoors. suitable for lawyer, businessmen, students, travelers and amateur archivists. Scan your important files and save them immediately, no struggling in finding a printing shop, keep it confidential.

For PDFs with visible horizontal or vertical rules, the reading-order stage can use those rules as structural signals when the PDF backend exposes their geometry. This behavior is enabled by default. Advanced options document how to disable it in Python or through the CLI flag --no-reading-order-separators. See advanced options and the pipeline options reference.

Review pages where order is ambiguous

  • Multi-column layouts, where lines from adjacent columns can be interleaved.
  • Ruled sections, where visible lines may affect how content is grouped.
  • Headers and footers, which should be distinguished from the main text.
  • Transitions between tables and surrounding paragraphs or captions.

Inspect representative pages in the JSON hierarchy or rendered Markdown and compare them with the PDF. Docling documents the representation and controls, but does not claim that every layout will be read in the intended order without review.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Sale
Epson Workforce ES-50 Compact & Lightweight Mobile Document Scanner
  • PORTABLE SCANNER FOR USE ON-THE-GO — The fastest and lightest mobile single-sheet-fed compact document scanner in its class¹
  • QUICK DOCUMENT SCANNING ― This Epson ultra-fast scanner scans a single page as quickly as 5.5 seconds²; Windows and Mac compatible
  • VERSATILE PAPER HANDLING ― Portable scanner scans documents up to 8.5 x 72 in; Also easily digitizes receipts and ID cards to make accounting, bookkeeping, and organizing simpler
  • INTUITIVE, HIGH-SPEED SOFTWARE — Epson ScanSmart Software³ is a smart tool allowing you to easily scan, review, and save; Stay organized easily with the help of this Epson scanner
  • EASY SETUP — USB-powered connect to your computer for quick and simple scanning; No batteries or external power supply required to operate portable document scanner; Standard Connectivity: USB 2.0
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Use OCR for image-only PDFs

OCR is relevant when a PDF contains scanned pages or images without a usable text layer. The documented full-page example is:

docling convert scan.pdf --ocr-mode full_page --to md

OCR adds processing time. If a PDF was generated digitally and already has an embedded text layer, check whether OCR is needed rather than enabling it automatically. The pipeline options reference describes OCR for scanned or image-based documents; the homepage shows the full-page command example.

Choose settings based on the PDF

  • Output purpose: Markdown for a readable deliverable; JSON when downstream software needs structured elements and reading order.
  • Table complexity: Use the faster approximate mode where suitable; consider the accurate TableFormer mode for complex or merged-cell tables, then verify the output.
  • Source condition: Use OCR for scanned or image-based pages; first check digital PDFs for an existing text layer.
  • Page structure: Keep reading-order separators enabled when visible rules may help, or disable them if your workflow requires that behavior.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.