Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Tesseract is an open-source optical character recognition (OCR) engine that extracts text from images. You can run it from a command line or call it through an API; it does not include a built-in graphical interface. To use it, install the engine and the trained data files for the language or script you want to recognize.
What Tesseract does—and what it doesn’t
Tesseract analyzes image content and produces recognized text or structured OCR output. It is a recognition engine, not a complete desktop scanning application: the official manual documents command-line and API use and notes that Tesseract has no built-in GUI. The manual describes the project as open source under the Apache 2.0 license. Tesseract user manual
The manual covers Tesseract 5.x and identifies version 5 as the current stable major version. It does not establish the latest minor or patch release, so consult the project’s release information for an exact version before choosing or documenting one.
Install the engine and the language data
Installation has two parts: the Tesseract program and the relevant traineddata files. The engine alone may not have the language or script support your images require. Package names, available language data, and installation locations vary by operating system and Linux distribution. Follow the instructions for your platform, then make sure the traineddata files are in a tessdata directory that Tesseract can find. Official installation guidance
#1 Best Overall
- OUR MOST ADVANCED SCANSNAP. Large touchscreen, fast 45ppm double-sided scanning, 100-sheet document feeder, Wi-Fi and USB connectivity, automatic optimizations, and support for cloud services. Upgraded replacement for the discontinued iX1600
- CUSTOMIZABLE. SHARABLE. Select personalized profiles from the touchscreen. Send to PC, Mac, mobile devices, and clouds. QUICK MENU lets you quickly scan-drag-drop to your favorite computer apps
- STABLE WIRELESS OR USB CONNECTION. Built-in Wi-Fi 6 for the fastest and most secure scanning. Connect to smart devices or cloud services without a computer. USB-C connection also available
- PHOTO AND DOCUMENT ORGANIZATION MADE EFFORTLESS. Easily manage, edit, and use scanned data from documents, receipts, photos, and business cards. Automatically optimize, name, and sort files
- AVOIDS PAPER JAMS AND DAMAGE. Features a brake roller system to feed paper smoothly, a multi-feed sensor that detects pages stuck together, and skew detection to prevent paper damage and data loss
There is no single installation command that applies to every platform. Linux distributions commonly provide packages for the engine and language data; other platform instructions use package managers, installers, or manual installation of .traineddata files.
Run OCR from the command line
The basic form is:
tesseract imagename outputbase
Here, imagename is the input image and outputbase is the output filename stem. The documented default is English with page segmentation mode 3. Specify a language explicitly with -l when needed:
Rank #2
- Design and Speed: Work with Windows XP/7/8/10/11 AND macOS 10.13 or later. Not compatible with Android and iOS. Designed for A3&A4(11.69*16.53 & 8.27*11.75 inch) document, any objects smaller than A3 size can be scanned with Ultra-fast scanning speed, about 1 second per page. Perfect device to scan FLAT papers
- USB Document Camera & Scanner: Work as both a document camera for remote teaching&learning compatible with ZOOM; Goole Meet and a document scanner to scan papers and convert/OCR files. OCR supports 180+ languages for text recognition. Please note that Thai, Hebrew, and Arabic are currently not supported. If you need the complete OCR language support list, please feel free to contact us for more details
- Patented Flattening Curved Book Page Technology: Shine Ultra applies CZUR’s patented technology to flatten the curved surface after pixel transformation to flattening of the book page (Only suitable for thinner books, ET series is recommended for thicker books)
- High Resolution & AI Tech: CMOS 13MP (4160*3120, A4≈340 AND A3≈245 DPI) camera. Smart Paging and Auto Cropping; Combine Sides; Stamp Mode; and Multiple Color Modes
- Height Adjustable & Portable: 2-level height adjustable neck. 90 degree foldable and lightweight 4 lbs with foot pedal for convenient operation
tesseract input.tiff output -l eng
Language codes must correspond to installed, discoverable traineddata. You can request multiple languages together by joining their codes with a plus sign; the installation guide’s example is:
tesseract input.tiff output -l eng+deu
For an explicit LSTM run, the command-line guide gives this form:
Rank #3
- FAST SPEEDS - Scans color and black and white documents a blazing speed up to 16ppm (1). Color scanning won’t slow you down as the color scan speed is the same as the black and white scan speed.
- ULTRA COMPACT – At less than 1 foot in length and only about 1. 5lbs in weight you can fit this device virtually anywhere (a bag, a purse, even a pocket).
- READY WHENEVER YOU ARE – The DS-640 mobile scanner is powered via an included micro USB 3. 0 cable allowing you to use it even where there is no outlet available. Plug it into you PC or laptop and you are ready to scan.
- WORKS YOUR WAY – Use the Brother free iPrint&Scan desktop app for scanning to multiple “Scan-to” destinations like PC, Network, cloud services, Email and OCR. (2) Supports Windows, Mac and Linux and TWAIN/WIA for PC/ICA for Mac/SANE drivers. (3)
- OPTIMIZE IMAGES AND TEXT – Automatic color detection/adjustment, image rotation (PC only), bleed through prevention/background removal, text enhancement, color drop to enhance scans. Software suite includes document management and OCR software. (4)
tesseract input.tiff output --oem 1 -l eng
The right settings depend on the image’s quality, page layout, language, and script. Tesseract’s documentation does not guarantee a particular accuracy for an untested input, so check results on representative pages before relying on them.
Choose an OCR engine mode and matching models
Tesseract 4.0 introduced an engine based on LSTM neural networks. In Tesseract 5, the documented engine mode options include --oem 1 for LSTM and --oem 0 for the legacy engine. Command-line usage guide
Rank #4
- FITS SMALL SPACES AND STAYS OUT OF THE WAY. Innovative space-saving design to free up desk space, even when it's being used
- SCAN DOCUMENTS, PHOTOS, CARDS, AND MORE. Handles most document types, including thick items and plastic cards. Exclusive QUICK MENU lets you quickly scan-drag-drop to your favorite computer apps
- GREAT IMAGES EVERY TIME, NO EXPERIENCE REQUIRED. A single touch starts fast, up to 30ppm duplex scanning with automatic de-skew, color optimization, and blank page removal for outstanding results without driver setup
- SCAN WHERE YOU WANT, WHEN YOU WANT. Connect with USB or Wi-Fi. Send to Mac, PC, mobile devices, and cloud services. Scan to Chromebook using the mobile app. Can be used without a computer
- PHOTO AND DOCUMENT ORGANIZATION MADE EFFORTLESS. ScanSnap Home all-in-one software brings together all your favorite functions. Easily manage, edit, and use scanned data from documents, receipts, business cards, photos, and more
- LSTM: Use
--oem 1with LSTM traineddata. The officialtessdata_bestandtessdata_fastmodel sets contain LSTM models only. - Legacy: Use
--oem 0when you need the legacy engine and have compatible legacy traineddata. Legacy models are included in thetessdatarepository’s model files.
Model availability and compatibility matter as much as the engine-mode flag: selecting a mode does not install its traineddata. Check which model set you installed and use a compatible engine mode.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Choose an output format for your workflow
Tesseract can produce several kinds of output. The useful choice depends on whether you need text alone, a document people can search, or OCR structure and coordinates. Official FAQ on output formats
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteBest Value
- FAST DOCUMENT SCANNING — Document scanner with feeder allows you to speed through stacks with a 50-sheet Auto Document Feeder (ADF); Efficient office scanner to help you scan more productively
- INTUITIVE, HIGH-SPEED SOFTWARE — Quickly scan with this desktop document scanner; Epson ScanSmart Software lets you easily preview scans, email files, upload to the cloud, and more; Plus, automatic file naming saves even more time
- SEAMLESS INTEGRATION — Easily incorporate your data into most document management software with the included TWAIN driver; Office document scanner integrates seamlessly with business workflows
- EASY SHARING — Duplex scanner allows you to scan straight to email or popular cloud storage2 services like Dropbox, Evernote, Google Drive, and OneDrive for simple storage and sharing
- SIMPLE FILE MANAGEMENT — Scanner allows the creation of searchable PDFs with Optical Character Recognition (OCR) and convert scans to editable Word or Excel files effortlessly; Designed for home and office document scanning
| Output | What it provides | Useful when |
|---|---|---|
| Plain text | Recognized text | You need text for review, editing, or another processing step. |
| Searchable PDF | The page image with a hidden searchable text layer | You want to preserve the scanned page’s appearance while making its text searchable. |
| hOCR | HTML-based OCR output that can include word coordinates | You need recognized words associated with positions on the page. |
| TSV | Tab-separated OCR output | You need structured results in a format suited to parsing or analysis. |
The command-line guide documents producing a searchable PDF with the pdf output configuration:
tesseract input.tiff output pdf
Output options and examples are described in the command-line usage guide.
Use an API or a third-party graphical interface
If you want to integrate OCR into software, the project provides a programmable API in addition to the command-line program. If you want point-and-click scanning, Tesseract itself does not supply that interface. The manual points users toward third-party GUIs, wrappers, and integrations, but their availability and maintenance vary; check the individual project before adopting one. Tesseract user manual
What to check when OCR results are poor
Start with the parts of the setup that determine whether Tesseract can interpret the input correctly:
Recommended Free Tools
Quick Recap
- Confirm that the traineddata for the image’s language or script is installed and discoverable.
- Check that the language codes passed to
-lmatch the installed files. - Verify that the selected
--oemmode is compatible with the traineddata model set. - Review whether the image quality and page layout are suitable for the default segmentation behavior; the basic command uses page segmentation mode 3.
- Compare output from representative pages before applying a configuration to a larger collection.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

