Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

No: the available evidence does not show that every developer has become a reviewer, or establish how AI-generated code changes review time or accuracy across the industry. Productivity studies report different results in different settings, and researchers have examined AI-assisted code review—but that is not the same as measuring the human reviewer’s workload or performance.

Does AI coding actually make developers more productive?

The answer depends on which developers, tasks, tools, and outcomes a study measures. Two prominent results point in different directions, but they are not estimates of the same thing.

Study Setting and participants Measured outcome and result
METR’s 2025 study, with a February 2026 update Experienced open-source developers worked on issues in their own repositories using AI tools available in early 2025. Tasks took 19% longer with AI, the study’s point estimate. METR’s 2026 update rounds this to a 20% slowdown in its opening summary and reports a confidence interval of +2% to +39%.
Management Science paper, published online February 27, 2026 Three randomized field experiments in ordinary business settings at Microsoft, Accenture, and an unnamed Fortune 100 company; 4,867 developers in the combined analysis. Developers using an AI code-completion assistant completed 26.08% more tasks on average (standard error 10.3%). The authors describe the experiments as noisy and the results as varying across them.

The first result concerns time to complete tasks for experienced open-source developers; the second concerns the number of completed tasks in workplace experiments. The studies also differ in work environment, tools, and study period. Neither result is a direct measure of code quality, reviewer accuracy, or review time, and the workplace average should not be read as a gain for every participant.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

METR’s February 2026 update says it believes developers are likely more sped up by AI in early 2026 than its early-2025 estimate suggested. It also warns that selection effects, reduced participation by developers unwilling to work without AI, and unreliable time reporting when participants used multiple agents make the later experiment weak evidence about the size of any change. These qualifications are specific to that later experiment; they do not turn the different studies into a single productivity estimate.

Are developers spending more time reviewing AI code?

The sources cited here do not establish how much additional human review time AI-generated code causes. They also do not provide a comparable cross-industry estimate of whether reviewers catch more or fewer defects in AI-assisted code than in code written without AI.

That is a distinct question from whether coding tools help developers complete tasks faster or increase task counts. To demonstrate a shift of work into review, evidence would need to measure that shift directly—for example, review time or review counts—rather than infer it from coding speed, tool adoption, or developers’ impressions of their work.

What do developers report about trusting AI-generated code?

The “Dear Diary” study, summarized by Microsoft Research and presented at ICSE-SEIP in April 2025, combined surveys, a randomized controlled trial, and a three-week diary study at a large multinational software company. Participants reported greater perceived usefulness and enjoyment after introduction and sustained use of coding tools, while their views about the trustworthiness of AI-generated code remained unchanged. Microsoft Research’s study summary reports that 84% of participants saw positive changes in daily work practices and 66% reported shifts in feelings about work.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Those percentages describe participant reports, not measurements of review accuracy, defect detection, or extra review hours. A change in how work feels or is organized can be important, but it cannot by itself show that reviewers are better—or worse—at evaluating generated code.

Has anyone tested AI code reviewers?

Yes. The 2024 ACM AIware paper “AI-Assisted Assessment of Coding Practices in Modern Code Review” describes AutoCommenter, an LLM-backed system for learning and applying coding best practices. The authors report implementing it for C++, Java, Python, and Go and evaluating it in a large industrial setting.

The paper distinguishes practices that can often be checked automatically, such as formatting rules, from guidance requiring more context or judgment. Nuanced conventions, exceptions in legacy code, and questions of clarity may depend on human knowledge. This work is evidence that AI-assisted review automation has been studied; it is not a test of the entire human-review role, nor evidence that reviewer workload or quality changed across the industry.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What would establish whether the reviewer is being tested?

A useful study would separate the work of producing code from the work of evaluating it, then measure both under clearly described conditions. To assess claims about the reviewer, look for evidence on:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Review effort: time spent reviewing, number of review rounds, and whether the figures distinguish generated from non-generated code.
  • Review performance: defects caught and missed, assessed against a credible reference for correctness.
  • Downstream results: defects that reach users or production, and maintenance consequences after the change is accepted.
  • Study context: participant experience, task type, AI tool and version, workplace or repository setting, and study dates.
  • Comparability: whether the AI-assisted and comparison groups handle similar tasks and whether the results are averages, ranges, or estimates with uncertainty.

Without those measures, a claim that AI has made developers spend more time reviewing—or has made review less effective—goes beyond what the cited evidence demonstrates. The studies support narrower conclusions about task productivity, work perceptions, and automated assessment of coding practices, each in its own setting.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.