iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Tom’s Guide reporter Amanda Caswell asked Claude to build a household expense tracker, then review it in three separate engineering roles. The passes surfaced different kinds of problems: functional bugs, accessibility and usability gaps, and a performance bottleneck. The experiment suggests a practical way to focus an AI coding assistant’s attention—but it does not show that role prompts alone make software correct or production-ready.
How the four-role experiment worked
In a report published August 30, 2026, Caswell described building a household expense tracker as a responsive Claude Artifact. The app was intended to let a household enter and delete expenses, assign categories, filter transactions, see totals, and view a visual category breakdown.
The initial prompt asked Claude to act as a senior full-stack engineer, plan the architecture, data structure, and user flow, and build the app. Caswell reported that the resulting React app included sample transactions, spending totals, category breakdowns, search, filtering, date sorting, and persistence. She then asked Claude to review the app through three narrower roles.
Free tools Windows power users keep installed
One-click scans. No signup required.
| Role or stage | Review objective | What the report says it surfaced | Evidence described |
|---|---|---|---|
| Full-stack engineer | Plan and build the app, including its architecture, data, and user flow. | A working tracker with transaction entry, totals, categories, search, filtering, sorting, and persistence. | Caswell’s account of the generated app; no underlying code is provided in the report materials described here. |
| Debugging engineer | Check functional behavior, input handling, calculations, persistence, data-loss risks, and edge cases before proposing fixes. | Problems with form structure, validation, empty-state calculations, immediate deletion, storage-error visibility, and missing automated tests. | Issues reported from Claude’s review, not a separate test report. |
| Frontend engineer | Review phone use, keyboard access, assistive technology, contrast, loading and error states, and destructive actions. | Missing focus indication, disconnected labels, placeholder-dependent search, validation announcements, contrast and small-screen search issues, and icon-only buttons without descriptive names. The pass also added a two-step deletion confirmation. | Findings and changes as described by Caswell; no independent accessibility audit is reported. |
| Performance engineer | Inspect rendering, calculations, sorting and filtering, storage, and memory against a target of at least 10,000 transactions. | A currency-formatting bottleneck and an optimization that reused a formatter. | Author-reported timings for the experiment; benchmark code and enough setup detail to reproduce them are not provided. |
What the debugging pass caught
The debugging prompt explicitly asked Claude to inspect functional behavior and data risks, and to explain issues and root causes before fixing them. According to Caswell, the review identified several weaknesses that could affect ordinary use:
#1 Best Overall
- Expense inputs were not organized in a proper form, and date and amount validation needed work.
- An empty tracker could incorrectly show Housing as its largest spending category.
- Deleting an expense happened immediately, without a confirmation step.
- Storage failures appeared only in the developer console, where a typical user would not see them.
- No automated tests had been created.
These are reported findings from Claude’s review, rather than independently verified defects. The missing tests matter: a code review can point to likely failure modes, but without tests or reproducible cases it does not establish that the proposed corrections work across relevant inputs.
Why the frontend pass found a different problem
The frontend prompt focused on how people interact with the tracker, especially on phones and with keyboards or assistive technology. Caswell reported that Claude found missing visible keyboard focus, labels that were not connected to their inputs, a search box relying on placeholder text instead of a proper label, and validation messages not configured for screen-reader announcement.
Rank #2
The pass also addressed contrast and small-screen search behavior, added descriptive names to icon-only buttons, and introduced a two-step confirmation for deletion. That confirmation is a useful example of the roles’ different focus: the debugging review had reportedly noticed that deletion was immediate, but the frontend review was the one that addressed the destructive-action experience.
Recommended Free Tools
Caswell also noted an inconsistency in the touch-target discussion. Claude described 44-by-44-pixel targets as a minimum, yet the main delete control was only 36 by 36 pixels. The report says that target met the smaller WCAG 2.2 AA target but did not meet the 44-pixel recommendation cited in the article. This is the report’s comparison, not an independent accessibility audit; it also illustrates why a general claim that an interface is accessible should be checked against the actual controls and criteria being applied.
Rank #3
What the performance numbers do—and do not—show
For the performance stage, the prompt set a target of at least 10,000 transactions and asked Claude to establish a baseline before optimizing. Caswell reported that formatting currency for 10,000 rows took 349 milliseconds in the initial version. Reusing one formatter reduced the reported time to 5.2 milliseconds, which the article characterized as a 67-fold improvement.
Those figures describe one author-reported experiment, not a general React or Claude performance guarantee. The report does not provide benchmark code or enough environment and setup detail for independent reproduction, so readers should not assume the same timings on other devices, implementations, or workloads.
Rank #4
The article also says Claude considered a household adding 30 to 50 transactions per month and concluded the original app would probably handle a decade of that usage without noticeable difficulty. That is Claude’s estimate as reported by Caswell, not the result of a decade-long test. The useful lesson is narrower: optimization should follow a workload and a measured baseline, rather than being added simply because a performance role was requested.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →What this workflow is useful for
Splitting a broad “make this production-ready” request into focused reviews can make the feedback easier to inspect. A debugging pass can concentrate on calculations and data handling; an accessibility pass can examine keyboard and assistive-technology interactions; and a performance pass can focus on measurements tied to an expected workload. Caswell’s account shows that the different prompts led to different reported findings, including a deletion issue that the earlier debugging pass had missed.
Best Value
That is a way to organize review, not a substitute for verification. Role labels do not prove that Claude found every defect, that a fix is correct, or that an app meets accessibility requirements. The reported absence of automated tests and the touch-target inconsistency are concrete reminders to verify important claims against the code and the relevant acceptance criteria.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.A practical way to adapt the approach
- Define the app and its real tasks. List what people must be able to do, what data the app stores, and what can go wrong. For a tracker, include entering, editing or deleting expenses, filtering records, and understanding totals.
- Ask for a build plan before implementation. Request an explanation of the architecture, data model, and user flow, then review that plan before asking for the app. This creates a concrete basis for later checks.
- Run a correctness review. Ask the assistant to inspect edge cases, validation, calculations, persistence, and possible data loss. Request issues and root causes separately from proposed fixes, then test the relevant cases.
- Run an interaction and accessibility review. Check visible keyboard focus, labels, screen-reader announcements, contrast, small-screen behavior, and the handling of destructive actions. Verify actual controls rather than relying on a blanket claim of accessibility.
- Set a workload and measure before optimizing. Specify a realistic volume and task, record a baseline under stated conditions, and only keep changes that measurably improve the behavior that matters.
- Turn findings into tests and user-visible checks. A review that identifies a risk is a starting point. Add tests or repeatable manual checks for important cases, and ensure failures such as unavailable storage are communicated to users rather than only logged for developers.
Caswell’s article warns that splitting the work into stages can consume Claude usage limits. It does not establish a particular plan, price, or amount of usage required, so the cost of this workflow will depend on the account and product limits in effect when it is used.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.

