iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
In Mercor’s comparison, Claude Opus 5 scored 100% on 20 attempts across four simplified month-end close scenarios, finishing each in under 10 minutes. The human comparison involved 12 licensed CPAs. That result shows strong performance on a narrow set of structured accounting exercises—not that AI can independently close a company’s books. A separate, broader accounting benchmark found that consistent end-to-end success remained rare.
What the CPA comparison actually tested
Mercor’s October 1, 2026 comparison put Claude Opus 5 and 12 licensed CPAs through four month-end close scenarios adapted from APEX-Accounting. The accountants averaged about five and a half years of experience. Each participant worked with company files, found figures, performed calculations, and returned results in tables. Mercor’s human-baseline report describes this as a comparison of unassisted humans and AI on medium-length, well-defined work.
Claude Opus 5 earned a perfect score across 20 attempts, with every attempt completed in less than 10 minutes. The result is striking, but its boundaries matter: the scenarios were simplified and designed around detail-oriented work, file search, calculations, and following close instructions. They did not represent every responsibility involved in closing a live company’s books.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsWhy the broader benchmark tells a different story
The human comparison should not be conflated with Mercor’s full APEX-Accounting benchmark. APEX covers 160 tasks across 10 fictional company worlds, authored by more than 40 accounting professionals with a median of 11 years’ experience. In Mercor’s July 2026 leaderboard, Claude Fable 5 led with 56.4% Mean Criteria@3. That score measures criteria met across evaluation runs; it is not a finding that the model completed 56.4% of full accounting workflows correctly. Mercor’s APEX report and the APEX-Accounting technical report describe the benchmark and its methodology.
#1 Best Overall
- Profitability calculations; cash flow function Calculates NPV and IRR for uneven cash flows
- Time-value-of-money and Amortization keys solve problems including: pension calculations, loans, mortgages, etc.
- Ideal calculator for students, managers and statisticians
- Built-in functionality : List-based one- and two-variable statistics with four regression options: linear, logarithmic, exponential and power
- The BA II Plus calculator is approved for use on the following professional exams: Chartered Financial Analyst exam. GARP Financial Risk Manager (FRM) exam. Certified Management Accountants exam
The more demanding measure is whether a model can complete a task correctly and repeatably. Mercor reports that the most consistent model solved only 2.6% of tasks correctly in all eight runs. In 58% of tasks, no model fully solved the task on any run. At least one model earned some credit on more than 95% of tasks, but partial credit is not a correct end-to-end close. Mercor’s own summary captures the distinction: “The headline result is not which model leads, but how rarely any model succeeds consistently.”
What these results do—and do not—establish
Promising territory: bounded, document-heavy tasks
The four-scenario comparison supports a limited conclusion: a capable model can perform very well on certain structured accounting exercises that involve finding details in files, following instructions, and calculating answers. Those are plausible areas to investigate for AI assistance, provided outputs are checked.
Rank #2
- Keys That Feel Right: Smooth, well-spaced keys with natural resistance allow you to move quickly and confidently—no re-learning or finger fatigue.
- Sharp, Color-Coded Printing: Prints 2.5 lines per second in black for positive and red for negative values—quiet, crisp, and easy to read at a glance.
- Big, Bright Display You Can Trust: The 12-digit fluorescent screen is clear from any angle, so totals are easy to catch without squinting or second-guessing.
- Designed for Speed and Comfort: Ergonomic key shapes follow your fingers’ natural motion—helping you type faster and make fewer mistakes.
- Built to Last, Easy to Maintain: Our heavy-duty design withstands daily use, featuring standard ribbons and paper rolls that are simple to replace.
Not demonstrated: autonomous month-end close
Neither study tested an AI system running a real company’s close from start to finish or signing off without human review. Mercor says the human comparison omitted client conversations, consultation with coworkers, clarifying questions, and accumulated business context. The APEX benchmark also excludes tax, audit, consolidation, multi-entity and multi-currency accounting, external reporting, and cases where an agent must ask for clarification. Its companies are fictional, and its results are benchmark results—not measured productivity gains in deployed businesses.
Mercor researcher Aden Barton cautions that the exercises tested “detail-oriented instruction following” rather than the whole accountant role, which includes less structured work such as client communication. That is Mercor’s interpretation of its study, not an independent consensus finding.
Rank #3
- Two-way Power Desk Calculator: Use solar power or battery power,In the case of sunlight or light, it can also be used without battery (Provide 2 AA batteries, only 1 needed).
- Optimized for Desk Use: The angled display offers a better viewing angle, especially when placed on a flat surface.
- Ergonomic Screen Tilt: Reduces neck strain with a user-friendly viewing angle, naturally aligning with your line of sight for a more comfortable experience.
- 10-Key Calculator with Large Buttons: Easy-to-use design follows computer keyboard layout.
- Desktop Basic Office Calculator:Perfect for daily use in offices, businesses, schools, retail stores, shopping centers, and home offices.
How to read speed, accuracy, and cost claims
For the simplified scenarios, Mercor reports that every Claude Opus 5 attempt took under 10 minutes. That is a task result in the study, not a forecast of real-world close time. Mercor also calculates cost at $0.21 per rubric criterion for Claude Opus 5 and $10.35 per criterion for unassisted accountants, using the US median accountant wage for the human-cost calculation. These are task-level figures, not estimates of total software ownership or deployment cost. The Decoder’s coverage also reports the comparison; Mercor is the primary source for the study figures.
A useful comparison of accounting AI claims should identify:
Rank #4
- Check your calculations thanks to the calculator's inbuilt serial impact roller printer that enables you to monitor your inputs and retains ongoing records. This two-color printer with a four-key memory prints red and black ink at up to 2.3 lines per second onto the included roll of paper.
- Printing calculator offers 12-digit LCD display for convenient viewing. 4-key memory keeps often-used figures accessible for faster calculations. Clock and calendar functions help maintain schedules.
- Easy-to-use solution for all of your basic math needs. Streamline financial calculations with currency conversion, tax calculation, and item counter functions. 1-year manufacturer limited warranty.
- Dimensions: 2.2"H x 6.4"W x 9.1"D. Package content: AC adapter, paper roll, user manual.
- Powered by any standard AC outlet, eliminating the need for expensive batteries. Decimal switch, rounding switch, percent, sign change, backspace, double zero, and grand total functions help you solve a variety of mathematical problems.
- Scope: whether the evaluation uses short, structured exercises or realistic accounting workflows.
- Completion standard: whether the score counts partial criteria or fully correct end-to-end tasks.
- Repeatability: how often the system succeeds across multiple runs, not just its best attempt.
- Conditions: the model snapshot, task set, tool access, and evaluation budget.
- Cost basis: cost per rubric criterion versus the full expense of implementing and supervising a system.
- Human context: whether the test includes clarification, communication, and business-specific judgment.
A score without those details is not enough to compare systems or infer that a company can safely hand over its close.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →What accounting teams can reasonably take from the studies
The findings support evaluating AI on specific, bounded tasks and designing review around the risks of each output. They do not establish that AI replaces accountants or can safely complete and sign off on a full close alone. Accountants remain central to supplying business context, resolving ambiguity, communicating with clients and colleagues, and verifying work—especially when a task does not fit a predefined instruction or document set.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

