What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

PDF extraction breaks retrieval-augmented generation (RAG) before the retriever ever runs. A text pass can return words in the wrong order, skip pages that are only images, flatten tables into noise, and drop figures entirely, and the chunks that reach the language model look fine until someone checks them. The practical fix is not a single “best parser.” It is to test each extraction stage separately on the documents you actually need to answer questions about.

This article explains where information goes missing, what the documented tools in this space can and cannot do, and how to compare extraction options on your own corpus. It does not report results from a specific custom pipeline, so every number below comes from the sources named in the text.

Where information is lost before retrieval

A RAG system can only answer from what ingestion preserves. Losses at this stage are easy to miss because the output is still readable text, and the failure only shows up later as a wrong or unsupported answer. Four failure points account for most of the trouble with scientific and administrative PDFs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Pages that contain no extractable text

Scanned pages and pages built from images have no text layer for a plain extractor to read. An extractor that returns an empty string for those pages is not reporting an error; it is reporting that it found nothing. A pipeline that only checks whether the job finished will pass that document as ingested.

#1 Best Overall
Canon Canoscan Lide 300 Scanner (PDF, AUTOSCAN, Copy, Send)
  • Scanner type: Document
  • Connectivity technology: USB
  • With Auto Scan Mode, the scanner automatically detects what you're scanning
  • Digitize documents and images

Reading order

Two-column layouts, sidebars, footnotes and captions are stored in the file in whatever order the producing software wrote them, which is often not the order a human reads them. Chunks built from that raw order can split a sentence across two unrelated paragraphs or attach a caption to the wrong figure. PyMuPDF documents extraction “in natural reading order” as a separate topic from basic text extraction, which is a signal that the two are not the same operation.

Tables

A table extracted as a stream of cell values loses its row and column relationships. A question such as “what was the threshold for category B?” may retrieve the right numbers with no way to tell which header they belong to. Tables need either structured output that keeps the grid or a textual description that preserves the relationships.

Figures

Figures carry information in charts, diagrams and images that text extraction does not read. Captions and surrounding paragraphs usually describe them, so a pipeline that stores only body text can answer questions about what a paper says about a figure while having no way to answer questions about the figure itself.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
Sale
Epson Workforce ES-50 Compact & Lightweight Mobile Document Scanner
  • PORTABLE SCANNER FOR USE ON-THE-GO — The fastest and lightest mobile single-sheet-fed compact document scanner in its class¹
  • QUICK DOCUMENT SCANNING ― This Epson ultra-fast scanner scans a single page as quickly as 5.5 seconds²; Windows and Mac compatible
  • VERSATILE PAPER HANDLING ― Portable scanner scans documents up to 8.5 x 72 in; Also easily digitizes receipts and ID cards to make accounting, bookkeeping, and organizing simpler
  • INTUITIVE, HIGH-SPEED SOFTWARE — Epson ScanSmart Software³ is a smart tool allowing you to easily scan, review, and save; Stay organized easily with the help of this Epson scanner
  • EASY SETUP — USB-powered connect to your computer for quick and simple scanning; No batteries or external power supply required to operate portable document scanner; Standard Connectivity: USB 2.0

Text extraction is not the same as OCR

Most PDF libraries treat extraction and optical character recognition (OCR) as two steps. PyMuPDF’s basic text documentation shows a plain text path using page.get_text(), and it separately tells users to apply OCR to pages whose text is image-based. The wording in that documentation is:

“If your document contains image based text content the use OCR on the page for subsequent text extraction:”

In PyMuPDF, that OCR step is performed with page.get_textpage_ocr() before the subsequent extraction. The practical consequence is that a pipeline must decide per page whether OCR is needed. Running OCR on every page adds cost and can introduce recognition errors into clean digital text; skipping it leaves scanned pages empty.

Rank #3
Sale
Brother DS-640 Compact Mobile Document Scanner, (Model: DS640)
  • FAST SPEEDS - Scans color and black and white documents a blazing speed up to 16ppm (1). Color scanning won’t slow you down as the color scan speed is the same as the black and white scan speed.
  • ULTRA COMPACT – At less than 1 foot in length and only about 1. 5lbs in weight you can fit this device virtually anywhere (a bag, a purse, even a pocket).
  • READY WHENEVER YOU ARE – The DS-640 mobile scanner is powered via an included micro USB 3. 0 cable allowing you to use it even where there is no outlet available. Plug it into you PC or laptop and you are ready to scan.
  • WORKS YOUR WAY – Use the Brother free iPrint&Scan desktop app for scanning to multiple “Scan-to” destinations like PC, Network, cloud services, Email and OCR. (2) Supports Windows, Mac and Linux and TWAIN/WIA for PC/ICA for Mac/SANE drivers. (3)
  • OPTIMIZE IMAGES AND TEXT – Automatic color detection/adjustment, image rotation (PC only), bleed through prevention/background removal, text enhancement, color drop to enhance scans. Software suite includes document management and OCR software. (4)

A stage-by-stage model for inspecting extraction

Treat ingestion as a sequence of checks, each of which can fail independently. The order below follows the path a document takes through the pipeline.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Classify each page. Determine whether it has an extractable text layer, is image-only, or is mixed. Record the count of pages in each class for every document.
  2. Apply OCR where required. Run OCR only on pages that need it, and keep the OCR output tagged so you can later measure its error rate separately from native text.
  3. Check reading order and hierarchy. Compare the extracted sequence against the rendered page for multi-column layouts, headings, footnotes and captions.
  4. Handle tables and figures explicitly. Confirm that each table keeps its header relationships and that each figure has a caption or description attached to it.
  5. Validate chunks and retrieval. Check whether each chunk is self-contained enough to answer a question, and whether the retriever returns the right chunk for questions you have written against the source.

Each stage produces its own evidence. A failure in step one cannot be fixed by a better chunker in step five.

Comparing the documented options

Three documented routes illustrate the main trade-offs. The table compares what each source states, and it leaves a cell as “not stated” where the cited source does not say.

Rank #4
Sale
Epson Workforce ES-400 II High-Speed Color Duplex Desktop Document Scanner
  • FAST DOCUMENT SCANNING — Document scanner with feeder allows you to speed through stacks with a 50-sheet Auto Document Feeder (ADF); Efficient office scanner to help you scan more productively
  • INTUITIVE, HIGH-SPEED SOFTWARE — Quickly scan with this desktop document scanner; Epson ScanSmart Software lets you easily preview scans, email files, upload to the cloud, and more; Plus, automatic file naming saves even more time
  • SEAMLESS INTEGRATION — Easily incorporate your data into most document management software with the included TWAIN driver; Office document scanner integrates seamlessly with business workflows
  • EASY SHARING — Duplex scanner allows you to scan straight to email or popular cloud storage2 services like Dropbox, Evernote, Google Drive, and OneDrive for simple storage and sharing
  • SIMPLE FILE MANAGEMENT — Scanner allows the creation of searchable PDFs with Optical Character Recognition (OCR) and convert scans to editable Word or Excel files effortlessly; Designed for home and office document scanning
Axis PyMuPDF core (get_text() and get_textpage_ocr()) PyMuPDF4LLM Adobe PDF Extract API
Where it runs Python library in your environment Python wrapper around PyMuPDF; hosting not stated Cloud API
Scanned pages Requires an explicit OCR call per page Not stated in the README Documented as handling native and scanned PDFs
Reading order Documented as a separate natural-reading-order extraction topic Combines text and tables in reading order Includes reading-order information in output
Tables Documented as a separate table-extraction topic Tables combined into Markdown output Documented as handling complex tables
Figures Documented as a separate image-extraction topic Not stated in the README Documented as including figures
Output Text and structured data you assemble yourself Markdown Structured JSON for downstream processing, or Markdown for LLM ingestion
Evidence available Official documentation Project README recommending it as a starting point for RAG Vendor documentation of capabilities

The sources do not include a controlled head-to-head comparison between PyMuPDF4LLM and Adobe’s API. Capability documentation tells you what a tool is designed to produce; it does not tell you how accurately it produces it on your documents. The vendor and project descriptions should be treated as starting claims to test, not as measured accuracy.

Local versus hosted processing

A local library keeps documents inside your environment and gives you full control over versions and OCR settings, but you own the work of handling tables and figures and of tuning the output. A managed API moves that work out of your code, but it introduces a dependency on a service, its availability and its terms, and it means sending document content to a third party. Confirm the current feature set, data handling and pricing in the vendor’s documentation before planning around a specific tier, since API features and commercial terms change.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the benchmark shows, and what it does not

A 2026 arXiv paper, From PDF to RAG-Ready: Evaluating Document Conversion Frameworks for Domain-Specific Question Answering, provides the most direct measured comparison among the sources cited here. It used a manually curated set of 50 questions over 36 Portuguese administrative documents, covering 1,706 pages and about 492,000 words. Answers were scored with an LLM-as-judge method.

Best Value
Sale
ScanSnap iX1300 Wireless or USB Double-Sided Color Document Scanner, Black
  • FITS SMALL SPACES AND STAYS OUT OF THE WAY. Innovative space-saving design to free up desk space, even when it's being used
  • SCAN DOCUMENTS, PHOTOS, CARDS, AND MORE. Handles most document types, including thick items and plastic cards. Exclusive QUICK MENU lets you quickly scan-drag-drop to your favorite computer apps
  • GREAT IMAGES EVERY TIME, NO EXPERIENCE REQUIRED. A single touch starts fast, up to 30ppm duplex scanning with automatic de-skew, color optimization, and blank page removal for outstanding results without driver setup
  • SCAN WHERE YOU WANT, WHEN YOU WANT. Connect with USB or Wi-Fi. Send to Mac, PC, mobile devices, and cloud services. Scan to Chromebook using the mobile app. Can be used without a computer
  • PHOTO AND DOCUMENT ORGANIZATION MADE EFFORTLESS. ScanSnap Home all-in-one software brings together all your favorite functions. Easily manage, edit, and use scanned data from documents, receipts, business cards, photos, and more
Configuration (as named in the paper) Reported score
Naïve PDFLoader 86.9%
Manually curated Markdown 97.1%
Docling with hierarchical splitting and image descriptions 94.1%

These scores describe one corpus, one set of pipeline configurations and one judging method. The manually curated Markdown result reflects human preparation of the text, so it shows the ceiling that careful preparation can reach rather than the output of an automated converter. The paper also reports that metadata enrichment and hierarchy-aware chunking contributed more to accuracy than converter choice alone. That finding argues for measuring chunking and metadata alongside extraction, not for choosing a converter in isolation.

How to evaluate extraction on your own documents

Build a small evaluation set from the documents your users will ask about. Use the following checklist.

  • Select a representative sample that includes native digital pages, scanned pages, multi-column layouts, dense tables and figure-heavy sections.
  • Write questions that can only be answered from specific tables, figures or footnotes, not only from body text.
  • Record the extracted output for each document and inspect it against the rendered page for reading order, table structure and missing content.
  • Run each candidate extractor and chunking setting on the same sample, and keep the configuration identical except for the variable under test.
  • Score retrieval first (is the correct chunk returned?) and answer quality second, so a generation error is not mistaken for an ingestion error.
  • Repeat the evaluation when you change the extractor version, OCR engine or chunk size, because those changes can shift results.

If you are building your own pipeline

A custom pipeline is worth documenting well enough that another engineer can reproduce your results. Record the following for each run:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • The extraction library and version, and the OCR engine and settings, if any.
  • The page classification rule and how many pages of each type the corpus contained.
  • How tables and figures are represented in the output, and how they are linked to their captions.
  • The chunking strategy, chunk size, overlap and any metadata attached to each chunk.
  • The evaluation set, the scoring method and the results broken down by document type, not only in aggregate.

Without those records, a score cannot be attributed to the extractor rather than to the chunker, the retriever or the corpus.

Quick Recap

Bestseller No. 1
Canon Canoscan Lide 300 Scanner (PDF, AUTOSCAN, Copy, Send)
Canon Canoscan Lide 300 Scanner (PDF, AUTOSCAN, Copy, Send)
Scanner type: Document; Connectivity technology: USB; With Auto Scan Mode, the scanner automatically detects what you're scanning
$75.00

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.