Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
PDFium and pypdf can produce different text from the same PDF because extracting text means reconstructing a reading experience from drawing instructions—not retrieving a dependable semantic transcript. Before sending extracted text to an LLM, check four areas: reading order, whitespace and layout, Unicode and ligatures, and scanned pages. These are practical mismatch categories, not a claim that every PDF or every version of the libraries exhibits all four.
Why the outputs can differ
PDFs primarily describe how content is drawn on a page. They do not reliably encode paragraph boundaries, table relationships, headers, or the order in which a person should read the page. As the pypdf documentation puts it, “PDF files don’t contain a semantic layer.” A text extractor therefore has to infer structure from positioned text and other page instructions, and different APIs or extraction modes can make different choices.
PDFium is the PDF engine. pypdfium2 is a Python wrapper around PDFium’s API; it is not the same thing as the separate pypdf library. The documentation describes mechanisms and limitations that make differences possible. It does not report a controlled, universal head-to-head result for every PDF. A specific file’s outcome depends on its contents, the library and engine versions, and the extraction settings.
Free tools Windows power users keep installed
One-click scans. No signup required.
Four mismatches to check
| Mismatch class | What the APIs and documentation establish | What to inspect before ingestion |
|---|---|---|
| Reading order | pypdf plain extraction follows text drawing commands in content-stream order; its documentation warns that the resulting order may change. PDFium exposes indexed page text, but that API shape is not a guarantee of natural reading order. | Compare extracted sequence with the rendered page, especially for columns, positioned labels, tables, footnotes, and floating figures. |
| Whitespace and layout | PDFium’s FPDFText_CountChars includes generated characters such as additional spaces and newlines. pypdf offers plain and layout extraction; layout mode reconstructs a fixed-width representation and has controls affecting vertical spacing and rotated text. |
Compare line breaks, spaces, blank lines, and the grouping of text. Neither source establishes that one library always preserves layout better or inserts more whitespace. |
| Unicode and ligatures | PDFium documents cases where a character cannot be converted to Unicode, and its text APIs have UCS-2 representation limits. pypdf identifies ligatures as an ambiguous case and shows post-processing to expand glyphs such as fi to fi. |
Compare code points as well as how strings look. Decide whether normalization is safe for the task rather than silently changing extracted text. |
| Scanned or image-only pages | pypdf is not OCR software and cannot extract text from images. A PDFium text API likewise cannot recover text that is absent from a page’s text layer. | If a page looks populated but yields little text, determine whether it has a text layer. Use OCR for image-only content and validate recognition results. |
1. Reading order: text can be present but sequenced badly
In pypdf’s plain extraction mode, text is located in the order drawing commands occur in the PDF content stream. That order can differ from the order a reader follows on the page. The project documentation explicitly cautions: “Do not rely on the order of text coming out of this function, as it will change if this function is made more sophisticated.” pypdf also offers an experimental layout mode, which represents a different extraction approach rather than a universal correction.
#1 Best Overall
- Scanner type: Document
- Connectivity technology: USB
- With Auto Scan Mode, the scanner automatically detects what you're scanning
- Digitize documents and images
PDFium’s page-text API exposes characters through an indexed text stream. The existence of that sequence does not establish that it will always match human reading order either. Inspect the output against a page rendering, particularly where multiple columns, sidebars, tables, footnotes, or floating figures compete for a place in the sequence. For an LLM, a misplaced heading or footnote can change what a nearby paragraph appears to mean.
2. Whitespace and layout: compare structure, not just words
Whitespace is part of extracted output, even when it is generated during extraction rather than represented as an ordinary printed character. PDFium documents that FPDFText_CountChars counts generated characters such as additional spaces and newlines. In pypdf, plain extraction and layout extraction produce different kinds of representations; layout mode reconstructs a fixed-width page and includes controls that affect vertical spacing and rotated text.
Rank #2
- PORTABLE SCANNER FOR USE ON-THE-GO — The fastest and lightest mobile single-sheet-fed compact document scanner in its class¹
- QUICK DOCUMENT SCANNING ― This Epson ultra-fast scanner scans a single page as quickly as 5.5 seconds²; Windows and Mac compatible
- VERSATILE PAPER HANDLING ― Portable scanner scans documents up to 8.5 x 72 in; Also easily digitizes receipts and ID cards to make accounting, bookkeeping, and organizing simpler
- INTUITIVE, HIGH-SPEED SOFTWARE — Epson ScanSmart Software³ is a smart tool allowing you to easily scan, review, and save; Stay organized easily with the help of this Epson scanner
- EASY SETUP — USB-powered connect to your computer for quick and simple scanning; No batteries or external power supply required to operate portable document scanner; Standard Connectivity: USB 2.0
Compare line endings, repeated spaces, blank lines, and whether text that belongs together stays together. A change in whitespace may be harmless for a bag-of-words task but consequential when an LLM must distinguish columns, list items, or table rows. The documentation supports checking these differences; it does not establish that either library always inserts more whitespace or preserves a page’s layout more faithfully.
3. Unicode and ligatures: visible similarity can hide string differences
A glyph that looks correct on the page may not map cleanly to a Unicode character during extraction. PDFium documents that FPDFText_GetUnicode can return zero when a character cannot be converted to Unicode. Its GetText API uses UCS-2 values and ignores characters without a UCS-2 representation. pypdf’s documentation calls out ligatures as an ambiguous extraction case and demonstrates post-processing substitutions such as expanding fi into fi.
Rank #3
- FAST SPEEDS - Scans color and black and white documents a blazing speed up to 16ppm (1). Color scanning won’t slow you down as the color scan speed is the same as the black and white scan speed.
- ULTRA COMPACT – At less than 1 foot in length and only about 1. 5lbs in weight you can fit this device virtually anywhere (a bag, a purse, even a pocket).
- READY WHENEVER YOU ARE – The DS-640 mobile scanner is powered via an included micro USB 3. 0 cable allowing you to use it even where there is no outlet available. Plug it into you PC or laptop and you are ready to scan.
- WORKS YOUR WAY – Use the Brother free iPrint&Scan desktop app for scanning to multiple “Scan-to” destinations like PC, Network, cloud services, Email and OCR. (2) Supports Windows, Mac and Linux and TWAIN/WIA for PC/ICA for Mac/SANE drivers. (3)
- OPTIMIZE IMAGES AND TEXT – Automatic color detection/adjustment, image rotation (PC only), bleed through prevention/background removal, text enhancement, color drop to enhance scans. Software suite includes document management and OCR software. (4)
When two outputs appear nearly identical, inspect their underlying code points: a ligature character and two ordinary letters may look similar while behaving differently in search, tokenization, or exact matching. pypdfium2 also warns that its range API is limited by UCS-2 and that the returned length can differ from the requested count in rare cases. Normalize text only if the downstream task permits the transformation, and preserve the original extraction if exact source fidelity may matter.
4. Scanned pages: extraction is not OCR
A scanned page may consist only of an image. pypdf states that it is not OCR software and cannot extract text from images; ordinary text extraction cannot create a text layer where none exists. If a visually populated page yields little or no text, check whether the PDF contains selectable text. If it is image-only, route it through OCR, then validate the result: OCR can misrecognize content, and a parser can also misread how an OCR text layer is represented.
Rank #4
- FAST DOCUMENT SCANNING — Document scanner with feeder allows you to speed through stacks with a 50-sheet Auto Document Feeder (ADF); Efficient office scanner to help you scan more productively
- INTUITIVE, HIGH-SPEED SOFTWARE — Quickly scan with this desktop document scanner; Epson ScanSmart Software lets you easily preview scans, email files, upload to the cloud, and more; Plus, automatic file naming saves even more time
- SEAMLESS INTEGRATION — Easily incorporate your data into most document management software with the included TWAIN driver; Office document scanner integrates seamlessly with business workflows
- EASY SHARING — Duplex scanner allows you to scan straight to email or popular cloud storage2 services like Dropbox, Evernote, Google Drive, and OneDrive for simple storage and sharing
- SIMPLE FILE MANAGEMENT — Scanner allows the creation of searchable PDFs with Optical Character Recognition (OCR) and convert scans to editable Word or Excel files effortlessly; Designed for home and office document scanning
A practical validation workflow for an LLM pipeline
- Pin the toolchain. Record the exact pypdf version, pypdfium2 version, and PDFium engine version used in the pipeline. Do not treat “PDFium” and “pypdfium2” as interchangeable package names.
- Record extraction settings. Keep the selected pypdf mode, layout options, pypdfium2 API calls, and any normalization or post-processing alongside the extracted text.
- Build a representative test set. Include multi-column pages, tables, rotated text, footnotes, unusual glyphs, and visually populated pages with little or no selectable text. Preserve the source PDFs and their provenance so a changed output can be traced to a changed file or toolchain.
- Compare against renderings. Check whether headings, columns, table cells, and footnotes appear in a sensible sequence. Review whitespace and Unicode where those features affect the task.
- Measure failures that matter to the task. Define what counts as a consequential error—such as a broken table row or a missing negation—rather than declaring one extractor “more accurate” without a target and reproducible evaluation.
- Route exceptions deliberately. Send image-only pages to OCR and validate the recognized text. Keep a review or fallback path for pages whose extraction does not meet the task’s requirements.
How to choose an extraction output
There is no evidence here for a universal winner. Choose based on the representation your application needs, then evaluate it on the PDFs and failure types your users actually encounter. If sequence matters, inspect reading order; if tables matter, inspect row and column relationships; if exact text matters, inspect Unicode and normalization; if pages are scans, provide an OCR path. Keep the rendered page available as a reference when extracted text is ambiguous.
Quick Recap
Best Value
- FITS SMALL SPACES AND STAYS OUT OF THE WAY. Innovative space-saving design to free up desk space, even when it's being used
- SCAN DOCUMENTS, PHOTOS, CARDS, AND MORE. Handles most document types, including thick items and plastic cards. Exclusive QUICK MENU lets you quickly scan-drag-drop to your favorite computer apps
- GREAT IMAGES EVERY TIME, NO EXPERIENCE REQUIRED. A single touch starts fast, up to 30ppm duplex scanning with automatic de-skew, color optimization, and blank page removal for outstanding results without driver setup
- SCAN WHERE YOU WANT, WHEN YOU WANT. Connect with USB or Wi-Fi. Send to Mac, PC, mobile devices, and cloud services. Scan to Chromebook using the mobile app. Can be used without a computer
- PHOTO AND DOCUMENT ORGANIZATION MADE EFFORTLESS. ScanSnap Home all-in-one software brings together all your favorite functions. Easily manage, edit, and use scanned data from documents, receipts, business cards, photos, and more
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

