Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The most dependable beginner workflow is Pillow + pytesseract + the Tesseract OCR engine. Pillow opens the JPG, pytesseract sends it to Tesseract, and image_to_string() returns recognized text.

Install the external Tesseract engine and the trained language data first, then install the Python packages in the same environment that runs your script:

python -m pip install Pillow pytesseract

Tesseract supports JPEG input through its image-reading layer. The example below handles ordinary printed English text; change lang when the matching trained data is installed.

1. Install the OCR engine and Python libraries

Install Tesseract separately

pytesseract is a Python wrapper; it does not contain the Tesseract executable. Install Tesseract for your operating system using the current instructions for that environment, and install the trained data for every language you plan to recognize. The requested language code must match an installed data file.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
ScanSnap iX2500 Wireless or USB High-Speed Document Scanner, Black
  • OUR MOST ADVANCED SCANSNAP. Large touchscreen, fast 45ppm double-sided scanning, 100-sheet document feeder, Wi-Fi and USB connectivity, automatic optimizations, and support for cloud services. Upgraded replacement for the discontinued iX1600
  • CUSTOMIZABLE. SHARABLE. Select personalized profiles from the touchscreen. Send to PC, Mac, mobile devices, and clouds. QUICK MENU lets you quickly scan-drag-drop to your favorite computer apps
  • STABLE WIRELESS OR USB CONNECTION. Built-in Wi-Fi 6 for the fastest and most secure scanning. Connect to smart devices or cloud services without a computer. USB-C connection also available
  • PHOTO AND DOCUMENT ORGANIZATION MADE EFFORTLESS. Easily manage, edit, and use scanned data from documents, receipts, photos, and business cards. Automatically optimize, name, and sort files
  • AVOIDS PAPER JAMS AND DAMAGE. Features a brake roller system to feed paper smoothly, a multi-feed sensor that detects pages stuck together, and skew detection to prevent paper damage and data loss

After installation, verify that the tesseract executable is available on your system PATH. If it is not, you can point pytesseract at the executable explicitly in Python.

Install Pillow and pytesseract

python -m pip install Pillow pytesseract

Use the same Python interpreter for installation and execution. A common cause of import errors is installing into one virtual environment and running the script with another.

2. Extract plain text from a JPG

Save this as extract_text.py beside scan.jpg:

from PIL import Image
import pytesseract

image = Image.open("scan.jpg")
text = pytesseract.image_to_string(image, lang="eng")
print(text)

Run it with:

python extract_text.py

image_to_string() returns a normal Python string. Tesseract may include line breaks and a trailing newline, so trim only when that is appropriate for your application:

clean_text = text.strip()
print(clean_text)

Use an explicit executable path when PATH is not configured

If pytesseract reports that it cannot find Tesseract, set tesseract_cmd before calling OCR. Replace the path with the executable location on your machine:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from PIL import Image
import pytesseract

pytesseract.pytesseract.tesseract_cmd = r"C:PathTotesseract.exe"
image = Image.open("scan.jpg")
print(pytesseract.image_to_string(image, lang="eng"))

On macOS or Linux, use the full path returned by your package manager or shell. Do not point to a directory; point to the Tesseract executable itself.

3. Select the correct language

The lang argument is a plus-separated list of installed languages. For example, eng+fra asks Tesseract to use English and French data:

Rank #2
Sale
CZUR Shine Ultra Smart Portable Document Scanner, Thin Book Scanner
  • Design and Speed: Work with Windows XP/7/8/10/11 AND macOS 10.13 or later. Not compatible with Android and iOS. Designed for A3&A4(11.69*16.53 & 8.27*11.75 inch) document, any objects smaller than A3 size can be scanned with Ultra-fast scanning speed, about 1 second per page. Perfect device to scan FLAT papers
  • USB Document Camera & Scanner: Work as both a document camera for remote teaching&learning compatible with ZOOM; Goole Meet and a document scanner to scan papers and convert/OCR files. OCR supports 180+ languages for text recognition. Please note that Thai, Hebrew, and Arabic are currently not supported. If you need the complete OCR language support list, please feel free to contact us for more details
  • Patented Flattening Curved Book Page Technology: Shine Ultra applies CZUR’s patented technology to flatten the curved surface after pixel transformation to flattening of the book page (Only suitable for thinner books, ET series is recommended for thicker books)
  • High Resolution & AI Tech: CMOS 13MP (4160*3120, A4≈340 AND A3≈245 DPI) camera. Smart Paging and Auto Cropping; Combine Sides; Stamp Mode; and Multiple Color Modes
  • Height Adjustable & Portable: 2-level height adjustable neck. 90 degree foldable and lightweight 4 lbs with foot pedal for convenient operation
text = pytesseract.image_to_string(image, lang="eng+fra")

If the requested trained data is missing, Tesseract raises an error or cannot recognize the intended language. Install the matching data and check that Tesseract can see its data directory before changing Python code.

4. Improve difficult JPGs without assuming one universal recipe

Tesseract performs image processing internally, but poor focus, low contrast, skew, heavy JPEG artifacts, and decorative layouts can still reduce recognition quality. Inspect the original image first, then test one change at a time on representative files.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Grayscale and contrast

from PIL import Image, ImageOps
import pytesseract

image = Image.open("scan.jpg")
gray = ImageOps.grayscale(image)
contrast = ImageOps.autocontrast(gray)
text = pytesseract.image_to_string(contrast, lang="eng")
print(text)

Autocontrast can help some scans and hurt others. Keep the original and compare the resulting text rather than applying it blindly.

Thresholding

from PIL import Image, ImageOps
import pytesseract

image = ImageOps.grayscale(Image.open("scan.jpg"))
thresholded = image.point(lambda pixel: 255 if pixel > 180 else 0)
print(pytesseract.image_to_string(thresholded, lang="eng"))

The threshold value is image-dependent. Test several values only when the source has uneven lighting or a faint background; a hard threshold can erase thin characters.

Choose a page-segmentation mode for the layout

Tesseract’s page-segmentation setting tells it what kind of layout to expect. A block of text, a single line, and a sparse label are different problems. Pass a configuration string and compare outputs:

config = "--psm 6"  # one uniform block of text
text = pytesseract.image_to_string(image, lang="eng", config=config)

There is no universally best mode. Try a mode that matches the actual page, and validate it on your own images, especially when columns, captions, or isolated labels are present.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Sale
Brother DS-640 Compact Mobile Document Scanner, (Model: DS640)
  • FAST SPEEDS - Scans color and black and white documents a blazing speed up to 16ppm (1). Color scanning won’t slow you down as the color scan speed is the same as the black and white scan speed.
  • ULTRA COMPACT – At less than 1 foot in length and only about 1. 5lbs in weight you can fit this device virtually anywhere (a bag, a purse, even a pocket).
  • READY WHENEVER YOU ARE – The DS-640 mobile scanner is powered via an included micro USB 3. 0 cable allowing you to use it even where there is no outlet available. Plug it into you PC or laptop and you are ready to scan.
  • WORKS YOUR WAY – Use the Brother free iPrint&Scan desktop app for scanning to multiple “Scan-to” destinations like PC, Network, cloud services, Email and OCR. (2) Supports Windows, Mac and Linux and TWAIN/WIA for PC/ICA for Mac/SANE drivers. (3)
  • OPTIMIZE IMAGES AND TEXT – Automatic color detection/adjustment, image rotation (PC only), bleed through prevention/background removal, text enhancement, color drop to enhance scans. Software suite includes document management and OCR software. (4)

5. Process more than one JPG

For a directory of images, keep failures attached to their filenames so one bad file does not hide the rest:

from pathlib import Path
from PIL import Image
import pytesseract

for path in sorted(Path("images").glob("*.jpg")):
    try:
        with Image.open(path) as image:
            text = pytesseract.image_to_string(image, lang="eng")
        print(f"n--- {path.name} ---n{text}")
    except Exception as exc:
        print(f"{path.name}: {exc}")

Use an additional glob for uppercase extensions if needed, such as *.JPG. Keep the source filename with the OCR result when creating searchable archives or downstream records.

6. Choose structured output when plain text is not enough

TSV data with word positions

TSV output includes recognized words and fields such as confidence and bounding-box coordinates. It is useful for highlighting text or selecting a region:

from PIL import Image
import pytesseract
from pytesseract import Output

image = Image.open("scan.jpg")
data = pytesseract.image_to_data(
    image,
    lang="eng",
    output_type=Output.DICT,
)

for i, word in enumerate(data["text"]):
    if word.strip():
        print({
            "text": word,
            "confidence": data["conf"][i],
            "left": data["left"][i],
            "top": data["top"][i],
            "width": data["width"][i],
            "height": data["height"][i],
        })

Confidence values are signals for review, not proof that a word is correct. Set your own acceptance rules for the document type.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

hOCR for layout-aware HTML

hocr_bytes = pytesseract.image_to_pdf_or_hocr(
    image,
    lang="eng",
    extension="hocr",
)
with open("scan.hocr", "wb") as output:
    output.write(hocr_bytes)

hOCR preserves positional information in an HTML-like format. Use it when a later process needs lines, words, or coordinates rather than a single string.

Searchable PDF

pdf_bytes = pytesseract.image_to_pdf_or_hocr(
    image,
    lang="eng",
    extension="pdf",
)
with open("scan-searchable.pdf", "wb") as output:
    output.write(pdf_bytes)

These outputs are documented alternatives to plain text. Pick the smallest representation that satisfies the next step in your pipeline.

Rank #4
Sale
ScanSnap iX1300 Wireless or USB Double-Sided Color Document Scanner, Black
  • FITS SMALL SPACES AND STAYS OUT OF THE WAY. Innovative space-saving design to free up desk space, even when it's being used
  • SCAN DOCUMENTS, PHOTOS, CARDS, AND MORE. Handles most document types, including thick items and plastic cards. Exclusive QUICK MENU lets you quickly scan-drag-drop to your favorite computer apps
  • GREAT IMAGES EVERY TIME, NO EXPERIENCE REQUIRED. A single touch starts fast, up to 30ppm duplex scanning with automatic de-skew, color optimization, and blank page removal for outstanding results without driver setup
  • SCAN WHERE YOU WANT, WHEN YOU WANT. Connect with USB or Wi-Fi. Send to Mac, PC, mobile devices, and cloud services. Scan to Chromebook using the mobile app. Can be used without a computer
  • PHOTO AND DOCUMENT ORGANIZATION MADE EFFORTLESS. ScanSnap Home all-in-one software brings together all your favorite functions. Easily manage, edit, and use scanned data from documents, receipts, business cards, photos, and more

7. Troubleshooting common failures

“No module named PIL” or “No module named pytesseract”

The packages are missing from the interpreter running the script. Install them with that interpreter:

python -m pip install Pillow pytesseract

In a virtual environment, activate it first and rerun the command.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

“Tesseract is not installed” or executable-not-found errors

The Python wrapper is present, but the external engine is absent or not on PATH. Install Tesseract, reopen your terminal so environment changes take effect, or assign pytesseract.pytesseract.tesseract_cmd to the full executable path.

Missing language-data errors

The lang value does not correspond to installed trained data, or Tesseract cannot locate that data directory. Install the requested language and verify the engine’s configured data location. Start with eng only after confirming English data is available.

The script cannot open the JPG

A .jpg suffix does not guarantee valid JPEG bytes. Confirm that the file is complete and that its actual encoding is supported. Try opening it with an image viewer or re-exporting it from the source application before debugging OCR settings.

The output is empty or inaccurate

  • Zoom in on the source and confirm the printed characters are genuinely legible.
  • Verify the language and page-segmentation assumptions.
  • Try grayscale, contrast adjustment, or thresholding on a copy.
  • Check that the text is printed rather than handwriting; this workflow is not a guarantee for handwriting recognition.
  • Use TSV or hOCR to inspect where Tesseract placed words when reading order looks wrong.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

8. Reliability, performance, and operational considerations

Image quality dominates results

OCR cannot recover characters that are missing, blurred, clipped, or hidden by compression artifacts. Preserve the highest-quality source available, avoid repeatedly re-saving JPGs, and inspect a sample of outputs before trusting an automated batch.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Sale
Epson Workforce ES-400 II High-Speed Color Duplex Desktop Document Scanner
  • FAST DOCUMENT SCANNING — Document scanner with feeder allows you to speed through stacks with a 50-sheet Auto Document Feeder (ADF); Efficient office scanner to help you scan more productively
  • INTUITIVE, HIGH-SPEED SOFTWARE — Quickly scan with this desktop document scanner; Epson ScanSmart Software lets you easily preview scans, email files, upload to the cloud, and more; Plus, automatic file naming saves even more time
  • SEAMLESS INTEGRATION — Easily incorporate your data into most document management software with the included TWAIN driver; Office document scanner integrates seamlessly with business workflows
  • EASY SHARING — Duplex scanner allows you to scan straight to email or popular cloud storage2 services like Dropbox, Evernote, Google Drive, and OneDrive for simple storage and sharing
  • SIMPLE FILE MANAGEMENT — Scanner allows the creation of searchable PDFs with Optical Character Recognition (OCR) and convert scans to editable Word or Excel files effortlessly; Designed for home and office document scanning

Memory and runtime

Large, high-resolution images contain more pixels for Tesseract to process. Resize only when the original is unnecessarily large and the characters remain legible; downscaling small text can remove critical detail. For repeated jobs, load one image at a time and write results incrementally instead of keeping an entire directory in memory.

Make failures observable

Record the filename, language, configuration, and exception for each job. Keep the original image beside the extracted result so a reviewer can compare questionable text with the source. For regulated or high-consequence records, add a human review step rather than treating OCR confidence as a correctness guarantee.

Or skip the browser setup

If the JPG must first be captured from a webpage, ScreenshotNeo can create a clean image before you run the OCR code above. It accepts consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers. It is a capture service, not an OCR engine, so pass the returned image to Tesseract afterward.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp

Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

See the ScreenshotNeo API documentation for capture parameters. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients. One thousand screenshots per month are free with no card; paid plans start at $5 for 3,000 shots, and every feature is included on every plan. Sign up for the free ScreenshotNeo plan.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

FAQ

Frequently Asked Questions

Can OCR reliably read text arranged in columns?

Not always. Reading order depends on the page-segmentation model and the image layout. Test a matching --psm configuration and inspect TSV or hOCR coordinates when column order matters.

Does a high confidence value prove that a word is correct?

No. Tesseract confidence is a review signal. For important records, compare flagged words with the JPG and use a human approval step.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.