To fine-tune a transformer for invoice recognition, first decide which fields the system must return, then prepare representative invoices with matching OCR text, page coordinates and field labels. LayoutLM-family models use OCR tokens and their positions to identify fields; Donut can instead generate a structured result directly from the page image without a separate OCR stage. In either case, evaluate on invoices from suppliers and layouts the model did not see during training. No single accuracy figure applies to arbitrary invoices.
What invoice recognition should return
Invoice recognition is document information extraction: the input is one or more invoice pages, and the output is a structured record. Agree on that record before collecting labels. A practical starting schema includes supplier and buyer names and addresses, invoice number, issue date, due date, supplier and buyer tax IDs, subtotal, tax, total, currency, and line items.
Define each field precisely. For example, decide whether an address is one combined string or separate street, city and postal-code fields; how absent values are represented; whether dates are normalized; and whether amounts retain the printed currency symbol or are stored as a numeric value plus a currency code. For line items, specify the expected columns, such as description, quantity, unit price, tax rate and line total. These are schema decisions, not details the model can resolve consistently on its own.
Keep the extracted value distinct from any normalized value when traceability matters. A review interface can show the source text and page region alongside the normalized result, making it easier to catch a mistaken digit or an incorrect date interpretation.
#1 Best Overall
- PORTABLE SCANNER FOR USE ON-THE-GO — The fastest and lightest mobile single-sheet-fed compact document scanner in its class¹
- QUICK DOCUMENT SCANNING ― This Epson ultra-fast scanner scans a single page as quickly as 5.5 seconds²; Windows and Mac compatible
- VERSATILE PAPER HANDLING ― Portable scanner scans documents up to 8.5 x 72 in; Also easily digitizes receipts and ID cards to make accounting, bookkeeping, and organizing simpler
- INTUITIVE, HIGH-SPEED SOFTWARE — Epson ScanSmart Software³ is a smart tool allowing you to easily scan, review, and save; Stay organized easily with the help of this Epson scanner
- EASY SETUP — USB-powered connect to your computer for quick and simple scanning; No batteries or external power supply required to operate portable document scanner; Standard Connectivity: USB 2.0
Choose between OCR-plus-layout and OCR-free extraction
| Approach | What goes into the model | Annotation and pipeline implications | Useful when |
|---|---|---|---|
| LayoutLM or LayoutLMv3 | OCR token text and each token’s two-dimensional page bounding box | Prepare OCR output, normalize coordinates to the model’s expected range, and align word-level labels with subword tokens. A token-classification head predicts labels over the text. | OCR text and coordinates are available and you want to classify page text using its position. |
| Donut | The document image, with a structured target representation | Fine-tune the model to generate the desired structure directly. The approach does not require a separate OCR output as its model input. | You prefer an image-to-structured-output pipeline rather than an OCR-plus-token-classification pipeline. |
Hugging Face describes LayoutLM as jointly modeling text and layout for scanned-document understanding and information extraction, and its documentation covers token classification and question answering workflows. Donut, introduced by NAVER AI Lab authors in 2021, is an OCR-free visual document-understanding transformer. “OCR-free” describes the model approach; it does not mean that the result is guaranteed to be correct or that output validation is unnecessary.
Compare the options against your own constraints: OCR dependency, need for bounding boxes, annotation format, languages, line-item table handling, performance on unseen layouts, latency, GPU memory, and how reliably the output can be made to conform to your schema. The available evidence does not establish a universal winner on these dimensions.
Rank #2
- FAST SPEEDS - Scans color and black and white documents a blazing speed up to 16ppm (1). Color scanning won’t slow you down as the color scan speed is the same as the black and white scan speed.
- ULTRA COMPACT – At less than 1 foot in length and only about 1. 5lbs in weight you can fit this device virtually anywhere (a bag, a purse, even a pocket).
- READY WHENEVER YOU ARE – The DS-640 mobile scanner is powered via an included micro USB 3. 0 cable allowing you to use it even where there is no outlet available. Plug it into you PC or laptop and you are ready to scan.
- WORKS YOUR WAY – Use the Brother free iPrint&Scan desktop app for scanning to multiple “Scan-to” destinations like PC, Network, cloud services, Email and OCR. (2) Supports Windows, Mac and Linux and TWAIN/WIA for PC/ICA for Mac/SANE drivers. (3)
- OPTIMIZE IMAGES AND TEXT – Automatic color detection/adjustment, image rotation (PC only), bleed through prevention/background removal, text enhancement, color drop to enhance scans. Software suite includes document management and OCR software. (4)
Prepare invoice data without leaking near-duplicates
Collect the range of documents you expect in production
Include variation in suppliers, invoice templates, languages, scan quality and page structure. If the deployed system will receive multi-page invoices, rotated scans or low-resolution images, include those cases in development and evaluation data rather than assuming a clean single-page sample represents them. Keep records of the source and any preprocessing applied to each document.
Split by supplier or template
Do not randomly split nearly identical invoices from the same supplier across training and test sets. Their repeated layout and wording can make the test look easier than the production task of handling a new template. Assign suppliers or template families to partitions so that near-duplicates do not cross from training into testing. A supplier-disjoint test set is a stronger check of generalization to unfamiliar layouts.
Recommended Free Tools
Rank #3
- Fast and Efficient: Scans both sides of a document at the same time, in color, at up to 45 pages per minute, with a 60 sheet automatic feeder, and one touch operation. Innovative Feeding System.
- Reliably Handles Many Different Document Types: Receipts, business cards, reports, contracts, long documents, thick or thin documents, and more. Monochrome LCD Display.
- Designed exclusively for the included Canon CaptureOnTouch software;TWAIN and ISIS drivers are not supported.
- Easy Setup: Simply connect to your computer using the supplied USB-C cable.
- Bundled Software: Includes easy-to-use Canon CaptureOnTouch scanning software.
Protect the source documents
Invoices can contain personal or commercially sensitive names, addresses, tax identifiers, dates and monetary amounts. Restrict access to raw images and annotations, retain only what the project needs, and consider redacted or synthetic examples for demonstrations. Include a privacy review in the project plan: research on document-understanding models has shown that sensitive fields can be reconstructed from some fine-tuning data.
Label fields in a way the model can learn
For a LayoutLM-family token classifier, run OCR and retain each token’s text and page bounding box. Convert bounding-box coordinates to the range expected by the selected model, then align the annotation with the model’s tokenization: one OCR word may become several subword tokens. Check this alignment carefully, because labels attached to the wrong tokens can undermine training even when the OCR appears correct.
Rank #4
- OUR MOST ADVANCED SCANSNAP. Large touchscreen, fast 45ppm double-sided scanning, 100-sheet document feeder, Wi-Fi and USB connectivity, automatic optimizations, and support for cloud services. Upgraded replacement for the discontinued iX1600
- CUSTOMIZABLE. SHARABLE. Select personalized profiles from the touchscreen. Send to PC, Mac, mobile devices, and clouds. QUICK MENU lets you quickly scan-drag-drop to your favorite computer apps
- STABLE WIRELESS OR USB CONNECTION. Built-in Wi-Fi 6 for the fastest and most secure scanning. Connect to smart devices or cloud services without a computer. USB-C connection also available
- PHOTO AND DOCUMENT ORGANIZATION MADE EFFORTLESS. Easily manage, edit, and use scanned data from documents, receipts, photos, and business cards. Automatically optimize, name, and sort files
- AVOIDS PAPER JAMS AND DAMAGE. Features a brake roller system to feed paper smoothly, a multi-feed sensor that detects pages stuck together, and skew detection to prevent paper damage and data loss
Annotate the text that expresses each field and decide how to handle repeated or ambiguous candidates. An invoice may show more than one date or amount; labels need to distinguish, for example, the invoice issue date from a due date and a grand total from a subtotal. Preserve page references and coordinates so predictions can be inspected against the document.
For Donut, pair each document image with a consistent structured target. Use the same field names, ordering and representation across examples, including a clear convention for missing fields and line-item arrays. Inconsistent target formats teach the model inconsistent output behavior.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Best Value
- FAST DOCUMENT SCANNING — Document scanner with feeder allows you to speed through stacks with a 50-sheet Auto Document Feeder (ADF); Efficient office scanner to help you scan more productively
- INTUITIVE, HIGH-SPEED SOFTWARE — Quickly scan with this desktop document scanner; Epson ScanSmart Software lets you easily preview scans, email files, upload to the cloud, and more; Plus, automatic file naming saves even more time
- SEAMLESS INTEGRATION — Easily incorporate your data into most document management software with the included TWAIN driver; Office document scanner integrates seamlessly with business workflows
- EASY SHARING — Duplex scanner allows you to scan straight to email or popular cloud storage2 services like Dropbox, Evernote, Google Drive, and OneDrive for simple storage and sharing
- SIMPLE FILE MANAGEMENT — Scanner allows the creation of searchable PDFs with Optical Character Recognition (OCR) and convert scans to editable Word or Excel files effortlessly; Designed for home and office document scanning
The University of Lisbon repository’s 2022 record describes an invoice study using 813 invoice pictures, annotated for company, address, date, document number, buyer and seller tax numbers, total amount and tax amount. That is an example of a field set used in a study, not a universal schema or a guarantee that the same volume of data is sufficient for another invoice population.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Fine-tune and validate in a controlled workflow
- Define the output contract. Write down required fields, formats, missing-value behavior and line-item structure before annotation begins.
- Build representative data partitions. Collect invoices across expected variation and reserve supplier- or template-disjoint documents for validation and final testing.
- Prepare model-specific inputs. For LayoutLM-family models, extract OCR tokens and normalized page boxes. For Donut, prepare page images and consistent structured targets.
- Annotate and audit. Label the chosen fields, inspect text-to-token alignment for LayoutLM, and check consistency of targets for Donut. Review examples with repeated dates, totals and tax values.
- Fine-tune the appropriate output head or sequence target. Use token classification for field labeling with LayoutLM or LayoutLMv3; fine-tune Donut to generate the target representation from images.
- Evaluate by field and failure mode. Score the supplier-disjoint set, inspect errors by field, layout, language, image quality and OCR failure, and improve the weakest part of the pipeline rather than relying on a single aggregate score.
- Test the deployed path. Validate the complete flow from PDF page conversion or image input through extraction, parsing and output checks, because preprocessing and formatting failures can occur outside the model itself.
Measure accuracy for the documents that matter
Report field-level precision, recall and F1 so teams can see whether a model misses values, predicts incorrect ones, or does both. Add exact-match checks where the full normalized value must be correct. For dates, currencies, totals and tax amounts, define normalization and any numeric tolerance before scoring; otherwise, harmless formatting differences and genuine extraction errors can be mixed together.
If line items matter, report a separate table or line-item metric. A model that identifies the invoice-level total correctly may still omit rows, merge descriptions or attach a quantity to the wrong item. Inspect those errors separately from header-field performance.
Hugging Face’s 2023 documentation snapshot lists FUNSD as 199 annotated forms with more than 30,000 words, SROIE as 626 training and 347 test receipt images, and RVL-CDIP as 400,000 document images across 16 classes. These figures describe datasets, not a forecast of invoice extraction accuracy. Public form, receipt and document-classification results do not establish performance on a company’s unseen invoice templates.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Quick Recap
Common failure modes and practical checks
- OCR mistakes: A wrong or missing OCR token can lead a text-and-layout model astray. Review errors against the original page and separate OCR failures from classification mistakes.
- Confused candidates: An invoice can contain several dates, identifiers and amounts. Check whether labels distinguish the intended field rather than simply capturing a plausible-looking value.
- Unseen layouts: A model can learn recurring positions and wording. Track performance by supplier or template, especially on the held-out group.
- Line-item structure: Rows and columns may wrap or shift across pages. Evaluate row association and cell completeness, not only whether individual values appear somewhere in the output.
- Invalid or inconsistent output: Check generated or parsed results against the schema, including required fields, data types and missing values. A syntactically valid result can still contain a wrong extraction.
- Privacy exposure: Limit access to training documents and check how raw inputs, annotations and model artifacts are retained and shared.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

