Docling converts varied files—such as PDFs, Office documents, web pages and images—into a shared structured representation called a DoclingDocument, which you can then export as Markdown, JSON, table files or chunks for a retrieval-augmented generation (RAG) workflow. The practical benefit is a consistent conversion pipeline; the important caveat is that extracted content, especially from scans and complex tables, still needs checking against the original.
What Docling does
Docling is an open-source document-processing toolkit available as a Python package, API and command-line interface. Its project overview describes a pipeline that parses diverse formats, including PDFs, into a common DoclingDocument representation. That gives downstream tasks a shared structure rather than a separate parser result for every file type. See the Docling project overview and the supported formats reference.
That intermediate representation can be exported in forms suited to different jobs: Markdown for reading or editing, JSON for structured processing, CSV or HTML for individual tables, and JSONL chunks for RAG pipelines. Docling is therefore best understood as a conversion and preparation step—not as a guarantee that every source document will be reconstructed perfectly or that the resulting data is ready to trust without review.
Which files can Docling process?
The official format reference lists PDFs; modern and legacy Office formats; OpenDocument; EPUB; Apple Pages and Keynote; Markdown and AsciiDoc; LaTeX; HTML, XHTML and MHTML; CSV; common raster images; audio and video; WebVTT; email; AFP; BoxNote; and specialized inputs such as DocLang, USPTO XML, JATS XML, XBRL XML, Docling JSON and EBCDIC. This is a list of supported format families, not a promise that every format works in every installation without added dependencies.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware match#1 Best Overall
- PORTABLE SCANNER FOR USE ON-THE-GO — The fastest and lightest mobile single-sheet-fed compact document scanner in its class¹
- QUICK DOCUMENT SCANNING ― This Epson ultra-fast scanner scans a single page as quickly as 5.5 seconds²; Windows and Mac compatible
- VERSATILE PAPER HANDLING ― Portable scanner scans documents up to 8.5 x 72 in; Also easily digitizes receipts and ID cards to make accounting, bookkeeping, and organizing simpler
- INTUITIVE, HIGH-SPEED SOFTWARE — Epson ScanSmart Software³ is a smart tool allowing you to easily scan, review, and save; Stay organized easily with the help of this Epson scanner
- EASY SETUP — USB-powered connect to your computer for quick and simple scanning; No batteries or external power supply required to operate portable document scanner; Standard Connectivity: USB 2.0
Check the format reference for the exact input you have. Some less common or legacy formats need optional extras or external software. For example, some legacy Office conversions require LibreOffice; audio/video processing requires the ASR extra, and video also requires ffmpeg. Match your setup to the files before converting a large collection.
Choose an output for the next task
| Output | Best suited to | What to know |
|---|---|---|
| Markdown | Reading, editing and straightforward text-based workflows | Human-readable structure; inspect layout-sensitive content after conversion. |
| JSON | Applications that need structured document content | Preserves the DoclingDocument serialization for downstream processing. |
| CSV or HTML table export | Working with an extracted table in a spreadsheet or web-oriented workflow | Export detected tables individually; compare the result with the source. |
| Chunked JSONL | Preparing document passages for a RAG pipeline | Chunk type and token options are configurable; chunking choices affect how context is represented. |
Docling also documents HTML, plain text, DocTags, DocLang XML and archives, WebVTT and LaTeX outputs. Image handling can use placeholders, embedded images or references, depending on the output and configuration. Consult the format reference and CLI documentation for supported output and option details.
How do I convert a PDF to Markdown?
The CLI reference documents conversion to Markdown and JSON. A basic command-line workflow is:
Rank #2
- FAST SPEEDS - Scans color and black and white documents a blazing speed up to 16ppm (1). Color scanning won’t slow you down as the color scan speed is the same as the black and white scan speed.
- ULTRA COMPACT – At less than 1 foot in length and only about 1. 5lbs in weight you can fit this device virtually anywhere (a bag, a purse, even a pocket).
- READY WHENEVER YOU ARE – The DS-640 mobile scanner is powered via an included micro USB 3. 0 cable allowing you to use it even where there is no outlet available. Plug it into you PC or laptop and you are ready to scan.
- WORKS YOUR WAY – Use the Brother free iPrint&Scan desktop app for scanning to multiple “Scan-to” destinations like PC, Network, cloud services, Email and OCR. (2) Supports Windows, Mac and Linux and TWAIN/WIA for PC/ICA for Mac/SANE drivers. (3)
- OPTIMIZE IMAGES AND TEXT – Automatic color detection/adjustment, image rotation (PC only), bleed through prevention/background removal, text enhancement, color drop to enhance scans. Software suite includes document management and OCR software. (4)
- Install Docling using the project’s current installation instructions and confirm that the PDF workflow’s required dependencies are available.
- Run
docling input.pdffrom a terminal, replacinginput.pdfwith your file’s path. The documented CLI conversion produces Markdown and JSON outputs by default; check the current CLI reference for output options and defaults. - Open the Markdown output and compare headings, reading order, page breaks and any important tables with the PDF. If you need structured fields rather than a readable document, use the JSON output as the basis for your next step.
The exact flags and defaults can change between releases, so use the current CLI reference for output selection, pipeline configuration and page-range options. For code-driven workflows, the v2 guide shows Python API examples for converting a single file or batches.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Can Docling read scanned PDFs?
Yes, scanned PDFs and images can be processed with OCR enabled. A scan contains page images rather than an existing selectable text layer, so OCR is needed to recognize its text. Docling exposes OCR settings for PDF and image workflows, including whether to force OCR over existing text and which language or engine to use.
Before processing a batch, identify whether the PDFs are born-digital, scanned or mixed. Select the appropriate OCR and pipeline settings, and use the CLI’s page-range options when you need to target particular pages. OCR output can be affected by language, scan quality and layout; check recognized text against the page, particularly for names, numbers and records where a transcription error would matter. See the CLI reference and project overview.
Rank #3
- FAST DOCUMENT SCANNING — Document scanner with feeder allows you to speed through stacks with a 50-sheet Auto Document Feeder (ADF); Efficient office scanner to help you scan more productively
- INTUITIVE, HIGH-SPEED SOFTWARE — Quickly scan with this desktop document scanner; Epson ScanSmart Software lets you easily preview scans, email files, upload to the cloud, and more; Plus, automatic file naming saves even more time
- SEAMLESS INTEGRATION — Easily incorporate your data into most document management software with the included TWAIN driver; Office document scanner integrates seamlessly with business workflows
- EASY SHARING — Duplex scanner allows you to scan straight to email or popular cloud storage2 services like Dropbox, Evernote, Google Drive, and OneDrive for simple storage and sharing
- SIMPLE FILE MANAGEMENT — Scanner allows the creation of searchable PDFs with Optical Character Recognition (OCR) and convert scans to editable Word or Excel files effortlessly; Designed for home and office document scanning
How can I extract tables from a PDF to CSV?
Docling’s PDF workflow can use table-structure extraction. After conversion, iterate over the detected tables, export each table to a DataFrame, and save it as CSV or HTML. The official table export example demonstrates this route.
- Convert the PDF with a pipeline and table settings appropriate to the document.
- Inspect the tables Docling detected in the resulting document.
- Export each table to a DataFrame, then save the DataFrame as CSV for spreadsheet or data-processing use; choose HTML if that fits your destination better.
- Compare rows, columns, headings and cell values with the original pages before using the exported data.
The example shows how to export tables; it is not a benchmark showing that all table layouts are reconstructed without errors. Merged cells, irregular layouts and scan quality are reasons to pay particular attention during review.
How do I get structured JSON from documents?
Convert the input to Docling’s JSON output when the next step needs the serialized DoclingDocument rather than a human-readable Markdown file. The JSON preserves that document representation, which can then be consumed by your own code or another processing stage. The v2 guide documents CLI and Python conversion examples; use the format reference to check the relevant input and output support.
Rank #4
- Scanner type: Document
- Connectivity technology: USB
- With Auto Scan Mode, the scanner automatically detects what you're scanning
- Digitize documents and images
JSON serialization is structured output, not proof that every field is correct. If you are extracting consequential values, retain a way to trace them to their source pages and validate them before loading them into a database or acting on them.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How should you prepare documents for RAG?
A typical RAG preparation flow converts documents, chooses a chunking strategy, and exports chunks as JSONL for ingestion by a retrieval system. Docling’s chunk output supports configurable chunk types and token options. Chunking and metadata choices matter: a passage that is too small may lose context, while one that is too large may be harder to retrieve usefully. Review the resulting chunks for reading order, headings and important relationships before indexing them.
One 2026 preprint, “From PDF to RAG-Ready,” compared four open-source PDF-to-Markdown frameworks across 19 pipeline configurations. On its manually curated benchmark of 50 questions from 36 Portuguese administrative documents (1,706 pages and about 492,000 words), it reported 94.1% automated accuracy for Docling with hierarchical splitting and image descriptions. Manually curated Markdown scored 97.1%, while a naïve PDFLoader baseline scored 86.9%. The authors discuss hierarchy-aware chunking and metadata enrichment as influential. Those results describe that corpus and comparison setup, not a general accuracy rate for Docling across languages, file types or pipelines. See the preprint.
Best Value
- OUR MOST ADVANCED SCANSNAP. Large touchscreen, fast 45ppm double-sided scanning, 100-sheet document feeder, Wi-Fi and USB connectivity, automatic optimizations, and support for cloud services. Upgraded replacement for the discontinued iX1600
- CUSTOMIZABLE. SHARABLE. Select personalized profiles from the touchscreen. Send to PC, Mac, mobile devices, and clouds. QUICK MENU lets you quickly scan-drag-drop to your favorite computer apps
- STABLE WIRELESS OR USB CONNECTION. Built-in Wi-Fi 6 for the fastest and most secure scanning. Connect to smart devices or cloud services without a computer. USB-C connection also available
- PHOTO AND DOCUMENT ORGANIZATION MADE EFFORTLESS. Easily manage, edit, and use scanned data from documents, receipts, photos, and business cards. Automatically optimize, name, and sort files
- AVOIDS PAPER JAMS AND DAMAGE. Features a brake roller system to feed paper smoothly, a multi-feed sensor that detects pages stuck together, and skew detection to prevent paper damage and data loss
Should conversion run locally or through a service?
The project documents both local execution and service-based conversion. Local processing may suit workflows where files should remain within a chosen environment, including sensitive or air-gapped settings. A remote service is another documented route. Decide where documents will be processed based on your deployment and data-handling needs, and verify the actual service or environment terms that apply to your use case. The availability of local execution is not itself a certification, compliance guarantee or assurance about every deployment.
A practical quality-control checklist
- Identify the source: distinguish text PDFs from scans and record any formats that require extra dependencies.
- Choose the destination first: use Markdown for reading, JSON for structured processing, table exports for tabular work, or chunked JSONL for a RAG workflow.
- Configure extraction deliberately: set OCR language and behavior for scans or mixed PDFs, and enable table extraction where tables matter.
- Review against the original: check reading order, key figures, table cells, OCR text and any fields that drive decisions.
- Validate the pipeline on representative files: document layouts and languages vary, so a successful conversion of one file does not establish quality for every file in a collection.
The available evaluation evidence is limited to a particular corpus and configured pipeline. It does not establish a single accuracy figure for every document type, language, scan quality or configuration. For production or high-stakes use, treat Docling as a way to create structured material for review and downstream work, not as a substitute for validating that material.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

