Recommended Free Tools
A PDF converter can jumble a two-column paper because a PDF records where text appears on a page, not necessarily the order a person should read it. To produce a sensible document, software must infer the page’s layout—often by detecting a whitespace gutter—and then order text within each region. A fixed “every page has two columns” rule is not enough: titles, tables, figures and other blocks may span or interrupt columns.
Why text extraction can alternate between columns
PDF text is positioned using coordinates and drawing instructions. The sequence in which text fragments are stored or emitted may differ from the visual reading order. A simple top-to-bottom sort across the entire page can therefore pick up a line from the left column, then one from the right, and repeat. It may also splice unrelated words together when lines from both columns share similar vertical positions.
OCR does not automatically solve this ordering problem. OCR recognizes text in an image; layout analysis decides how recognized text belongs in regions and in what sequence those regions should be read. Recognition and reading-order reconstruction are separate stages.
How a converter should find columns
Use page geometry to identify likely regions
A common geometric strategy is to collect text fragments and their bounding boxes, measure where text occupies the page horizontally, and look for a substantial interior whitespace band. That band may indicate the gutter between columns. Once likely regions are identified, the converter can order text top-to-bottom within the left region, then within the right, instead of sorting every line across the full page.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →#1 Best Overall
- Scanner type: Document
- Connectivity technology: USB
- With Auto Scan Mode, the scanner automatically detects what you're scanning
- Digitize documents and images
This is a useful implementation pattern, not a universal standard. Google Research’s Ray Smith described a different but related method: infer column layout from formatting tab stops, then apply that layout from the top down to impose structure and reading order (Hybrid Page Layout Analysis via Tab-Stop Detection, 2009).
Segment mixed layouts into bands
A gutter is evidence of a boundary, not proof that the whole page follows one two-column template. A full-width title or heading can sit above the columns; an abstract, table or figure may span them; and a reference list may use a different arrangement. A robust converter separates full-width blocks from column regions, orders each region appropriately, and places the blocks back at their vertical positions.
Rank #2
- PORTABLE SCANNER FOR USE ON-THE-GO — The fastest and lightest mobile single-sheet-fed compact document scanner in its class¹
- QUICK DOCUMENT SCANNING ― This Epson ultra-fast scanner scans a single page as quickly as 5.5 seconds²; Windows and Mac compatible
- VERSATILE PAPER HANDLING ― Portable scanner scans documents up to 8.5 x 72 in; Also easily digitizes receipts and ID cards to make accounting, bookkeeping, and organizing simpler
- INTUITIVE, HIGH-SPEED SOFTWARE — Epson ScanSmart Software³ is a smart tool allowing you to easily scan, review, and save; Stay organized easily with the help of this Epson scanner
- EASY SETUP — USB-powered connect to your computer for quick and simple scanning; No batteries or external power supply required to operate portable document scanner; Standard Connectivity: USB 2.0
Recursive whitespace segmentation, often described as XY-Cut, is one way to divide a page into regions. OpenDataLoader’s documentation describes first separating cross-layout elements such as full-width titles and headers, then segmenting the remaining content and restoring those elements in reading order (Reading Order & XY-Cut++). Other implementations use horizontal bands to distinguish spanning elements from column text. These are approaches, not guarantees that every page will be interpreted correctly.
Born-digital PDFs and scanned papers need different handling
Born-digital PDFs
When a paper contains selectable text, a converter can use character data, fonts and text positions to infer paragraphs and columns. That positional information can make layout recovery easier, but it does not guarantee that the PDF’s internal sequence matches human reading order.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsRank #3
- FAST SPEEDS - Scans color and black and white documents a blazing speed up to 16ppm (1). Color scanning won’t slow you down as the color scan speed is the same as the black and white scan speed.
- ULTRA COMPACT – At less than 1 foot in length and only about 1. 5lbs in weight you can fit this device virtually anywhere (a bag, a purse, even a pocket).
- READY WHENEVER YOU ARE – The DS-640 mobile scanner is powered via an included micro USB 3. 0 cable allowing you to use it even where there is no outlet available. Plug it into you PC or laptop and you are ready to scan.
- WORKS YOUR WAY – Use the Brother free iPrint&Scan desktop app for scanning to multiple “Scan-to” destinations like PC, Network, cloud services, Email and OCR. (2) Supports Windows, Mac and Linux and TWAIN/WIA for PC/ICA for Mac/SANE drivers. (3)
- OPTIMIZE IMAGES AND TEXT – Automatic color detection/adjustment, image rotation (PC only), bleed through prevention/background removal, text enhancement, color drop to enhance scans. Software suite includes document management and OCR software. (4)
Scanned PDFs
A scan is an image of a page rather than a reliable text layer, so OCR must first recognize characters and estimate their locations. Skew, low contrast, noise and narrow gutters can reduce recognition quality or make region boundaries harder to determine. Some PDFs mix scanned pages with digital text, so the appropriate input path may differ from page to page. The all2md 1.14.0 documentation discusses this distinction and column-detection controls (PDF converter documentation).
Even accurate OCR can yield the wrong sequence if the layout stage fails. When a result looks scrambled, check both whether the words were recognized correctly and whether the regions were ordered correctly.
Rank #4
- FAST DOCUMENT SCANNING — Document scanner with feeder allows you to speed through stacks with a 50-sheet Auto Document Feeder (ADF); Efficient office scanner to help you scan more productively
- INTUITIVE, HIGH-SPEED SOFTWARE — Quickly scan with this desktop document scanner; Epson ScanSmart Software lets you easily preview scans, email files, upload to the cloud, and more; Plus, automatic file naming saves even more time
- SEAMLESS INTEGRATION — Easily incorporate your data into most document management software with the included TWAIN driver; Office document scanner integrates seamlessly with business workflows
- EASY SHARING — Duplex scanner allows you to scan straight to email or popular cloud storage2 services like Dropbox, Evernote, Google Drive, and OneDrive for simple storage and sharing
- SIMPLE FILE MANAGEMENT — Scanner allows the creation of searchable PDFs with Optical Character Recognition (OCR) and convert scans to editable Word or Excel files effortlessly; Designed for home and office document scanning
Why a fixed two-column assumption breaks
- Full-width blocks: A title or section heading may span the page, so it should not be treated as part of either column.
- Interrupted columns: A table or figure can cross or divide the column flow, changing the order in which surrounding text should be read.
- Layout changes: Abstracts, body pages and references may use different arrangements, even within one paper.
- Ambiguous geometry: A narrow gutter, columns that begin at similar heights, or noisy scan coordinates can confuse simple thresholds.
- Recognition errors: Skew or poor image quality can distort text boxes and make an otherwise plausible gutter harder to identify.
Geometric rules are often lightweight and explainable, but their thresholds can fail on irregular layouts. Learned or semantic layout analysis can classify region types and reading order, but adds model and dependency considerations and still requires verification. No neutral, broadly representative accuracy comparison establishes one approach as best for all papers.
What to do when your converted text is out of order
- Check whether the PDF has selectable text. If it does, the issue may be layout reconstruction rather than OCR. If it is a scan, confirm that OCR was applied and inspect whether the recognized words themselves are correct.
- Look for layout controls in the converter. If available, enable column detection or set the page’s column count. Use a two-column override only for pages that actually have that layout; a paper with mixed page types may need per-page or per-region handling.
- Test representative page types. Compare the conversion with the visible first page, an ordinary body page, a page containing a figure or table, and the references. Check that headings precede their text, paragraphs do not alternate between unrelated columns, and reference entries remain intact.
- Use manual ordering when automatic detection fails. If the software exposes page regions or an order list, split a wrongly grouped two-column region and move the pieces into the intended sequence.
- Choose another extraction route if needed. Compare results on the same representative pages, including how each tool handles full-width blocks, figures, tables, references, manual correction, privacy, performance and licensing. A converter that works on a plain body page may still fail on the paper’s first page or references.
Manual reading-order repair in Adobe Acrobat Pro
Adobe documents a manual workflow for adjusting tagged content with the Reading Order tool. Its help says, “If a single highlighted region contains two columns of text or text that won’t flow normally, divide the region into parts that can be reordered.” Adobe also describes changing an item’s order in the Order panel or by dragging it on the page; this changes reading order without changing the visible appearance of the PDF (Reading Order tool for PDFs).
Best Value
- FITS SMALL SPACES AND STAYS OUT OF THE WAY. Innovative space-saving design to free up desk space, even when it's being used
- SCAN DOCUMENTS, PHOTOS, CARDS, AND MORE. Handles most document types, including thick items and plastic cards. Exclusive QUICK MENU lets you quickly scan-drag-drop to your favorite computer apps
- GREAT IMAGES EVERY TIME, NO EXPERIENCE REQUIRED. A single touch starts fast, up to 30ppm duplex scanning with automatic de-skew, color optimization, and blank page removal for outstanding results without driver setup
- SCAN WHERE YOU WANT, WHEN YOU WANT. Connect with USB or Wi-Fi. Send to Mac, PC, mobile devices, and cloud services. Scan to Chromebook using the mobile app. Can be used without a computer
- PHOTO AND DOCUMENT ORGANIZATION MADE EFFORTLESS. ScanSnap Home all-in-one software brings together all your favorite functions. Easily manage, edit, and use scanned data from documents, receipts, business cards, photos, and more
This is a repair workflow for tagged PDF content, not a one-click fix for every converter or text extractor. It is most relevant when you can edit the PDF’s reading order itself; an external conversion pipeline may require its own layout settings or manual correction.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

