What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

AI can make test-case generation dramatically cheaper in a narrowly defined workflow, but the strongest direct comparison is a proof of concept for synthetic survey data—not a general benchmark for software teams. A 2025 National Cancer Institute study estimated $381 and eight hours per manually generated case, versus $0.10 and 16.5 minutes or 3.75 minutes for two AI approaches. Those AI figures excluded setup and deployment of the generation framework, as well as the time needed to train a person to answer the survey manually.

What the direct cost comparison measured

The National Cancer Institute (NCI) project generated synthetic answers for three surveys in the CHARMS Rasopathy workflow. The aim was to run existing automated tests without using identified patient-level production data. Manually, testers traversed each survey, copied its questions into an input file, and created answers. The automated workflow extracted questions from survey JSON, used a persona and question dependencies to generate synthetic responses, and packaged the results for testing.

The study’s per-case estimates were:

Approach Estimated time per case Estimated cost per case What the estimate represents
Manual case generation 8 hours $381 NCI authors’ estimate based on an average automation tester salary of $99,000.
Azure OpenAI GPT-3.5 16.5 minutes $0.10 NCI authors’ estimate for the study workflow; the reported time included waiting between API calls.
Self-hosted AWS Flan T5-XL 3.75 minutes $0.10 NCI authors’ estimate for the study workflow; its self-hosted endpoint did not require the same API-call wait.

These are historical study configuration figures, not current cloud or software price quotes. The $381 manual estimate depends on the salary assumption, and the $0.10 AI figures do not include the cost of building and deploying the framework or training a person to answer the survey manually. The study generated 50 cases with each AI approach. Read the NCI study.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to interpret the headline savings

The NCI authors wrote: “Synthetic data generation is greater than 3,000X cheaper and greater than 120X faster than the manual test case generation process”. That statement describes their synthetic survey-data workflow and its stated per-case calculation; it is not a total-cost-of-ownership result or a promise for other teams. The excluded framework work, deployment, and manual training affect the comparison, and the manual labor estimate reflects the study’s salary assumption.

There is also a quality boundary to the cost result. Survey responses followed conditional paths, and the authors said covering every possible path was impractical. Their evaluation considered response completeness, complexity of text answers, demographic coverage, and clinical expert review. They noted demographic omissions in generated data, even though some categories were represented better than in manually created test data. Lower generation effort alone does not establish that cases are complete, realistic, or useful for finding defects.

Why other testing studies do not produce one savings rate

Published comparisons cover different jobs: authoring and maintaining executable scripts, designing tests from requirements, and executing existing manual tests. Their results can inform a team’s decision, but they should not be combined with the NCI synthetic-data figures as if they measured the same task.

Writing and evolving web test scripts

Leotta, Ricca, Marchetto, and Olianas compared an NLP-based approach with Selenium WebDriver and Selenium IDE across nine test suites on different web applications. Three junior testers or developers, each with roughly two to three years of end-to-end web-testing experience, took part. The study assessed initial development, how many scripts remained reusable after an application version change, time to evolve suites, and cumulative effort. The authors concluded that NLP-based automation appeared competitive for small-to-medium suites “such as those considered in our empirical study.” This is a lifecycle-cost perspective, not a general LLM savings percentage. Read the study.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Generating scripts from written test cases

A 2024 preliminary study looked at ChatGPT and GitHub Copilot producing web end-to-end test scripts from natural-language descriptions. It reported reduced development time when the input was clearly defined in Gherkin, while cautioning that testers need enough scripting skill to modify AI-produced code. The public repository record does not provide a numeric breakdown, so it cannot support a precise percentage saving. See the repository record.

Designing system tests from user stories

A 2025 public-sector study described a GPT-4 tool connected to Redmine and Squash TM. Analysts said it reduced effort, and the study reported the generated and manually designed tests had the same functional coverage. The accessible study page gives no quantified time or money comparison. Matching functional coverage is useful evidence about that workflow, but it does not establish equal coverage for other systems or a general cost reduction. Read the study.

Executing existing manual regression tests

Bauer, Frattini, and Alégroth evaluated Augmented Testing, a visual support layer for manual GUI regression work, with 13 professionals from six companies. Mean execution time across all tests was 1,222 seconds for manual GUI execution and 779 seconds with the assistance, a 36% reduction. Six of eight cases were faster; the two shortest slightly favored the baseline. This measures execution assistance, not the cost of creating test cases. Read the study.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to calculate your team’s real cost

Compare the cost of accepted, usable tests over the releases you expect to support—not just model charges or the time to produce a draft. A useful scenario calculation is:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Total effort or cost over a chosen period = initial setup + generation + human review and correction + maintenance + tool or cloud charges.

Track each component separately so the result reflects your workflow:

  • Initial setup: Include framework or prompt-workflow construction, data preparation, integrations, and staff onboarding. The NCI per-case figures excluded framework build and deployment; the manual estimate excluded training time.
  • Human review and correction: Record time to check expected results, repair errors, and assess domain-specific and boundary conditions. Script-generation results depend on testers being able to modify generated code.
  • Coverage and realism: Compare functional and branch coverage, boundary cases, realistic data, and defect detection—not raw case counts. The NCI study found demographic omissions; the public-sector study reported matching functional coverage for its specific system.
  • Maintenance: Measure which tests survive application changes and the effort to update those that break. The Leotta study explicitly included suite evolution and cumulative effort.
  • Execution: Keep the time spent running tests separate from authoring or generating them. Assisted manual execution can change regression time without reducing case-design effort.
  • Local labor and usage rates: Substitute your organization’s loaded labor cost and current tool or cloud rates for historical study assumptions.

The break-even point depends on scale: recurring generation and maintenance savings must exceed setup, review, and correction effort across the cases and releases your team actually needs. The cited studies do not provide a standardized cross-study benchmark, so measure the same workflow and quality criteria on both sides of your own comparison.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.