Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

AI coding tools have changed how some developers write and review code, but the evidence does not show that every developer has become a reviewer or that software teams are broadly getting worse. Studies measure different outcomes—from task completion and code quality in controlled exercises to developers’ perceptions—and do not yet establish whether AI-generated code increases total review work or worsens long-term production results.

What does the evidence actually show?

Studies of AI-assisted coding have produced results that depend on the developers, tasks, tools, and outcomes being measured. More tasks completed does not necessarily mean better software; a passing test suite does not establish low long-term defect rates; and a developer’s view of a tool does not measure reviewer workload.

Study Setting and measure Reported result What it does not establish
INFORMS / Management Science, online February 27, 2026 Three randomized field experiments across Microsoft, Accenture, and an anonymous Fortune 100 company; 4,867 developers. Outcome: completed tasks. AI-tool users completed 26.08% more tasks on average; standard error was 10.3%. Effects varied among the experiments, and less experienced developers showed higher adoption and larger gains. Whether the additional tasks were higher quality, created more review work, or produced better long-term software outcomes.
METR, July 12, 2025 Randomized study of 16 experienced developers performing 246 tasks in familiar, mature open-source projects, with tools available from February to June 2025. Outcome: task completion time. Participants estimated that AI would reduce their time by 20%; measured completion time was 19% longer. A general effect across developers or projects. The authors said experimental artifacts could not be entirely ruled out.
GitHub, November 18, 2024; page updated February 6, 2025 Vendor-published controlled study: experienced Python developers built API endpoints for a fictional restaurant-review server. There were 202 valid submissions in the first phase and 25 developers blind-reviewed qualifying submissions. Copilot-group submissions were 53.2% more likely to pass all 10 unit tests; reviewers found 13.6% more lines per readability error. GitHub also reported higher ratings for readability, reliability, maintainability, and conciseness, and a 5% higher likelihood of approval. Code quality in production systems, long-term defect rates, or review workload across organizations.
Microsoft Research, ICSE-SEIP 2025 Mixed-methods study at one large multinational software company, combining surveys, a randomized trial, and a three-week diary study. Developers increasingly viewed tools as useful and enjoyable, while views of generated-code trustworthiness remained unchanged. 84% reported positive changes in daily practices and 66% noted shifts in feelings about work. Whether changed perceptions corresponded to fewer defects, more reviewer hours, or a broad change in job roles.
Google Research, 2024 Industrial deployment of AutoCommenter, which assesses coding-language best practices for C++, Java, Python, and Go. The public abstract reports a measurable positive workflow impact. The abstract does not quantify reviewer hours, defect rates, or changes in reviewer roles.

These results should not be ranked as if they were head-to-head tests. The company field experiments and METR trial differ in population, codebase familiarity, task type, tool vintage, and setting. GitHub’s controlled exercise gives a defined quality and review measure, but its vendor-published result is not a production-system evaluation. Each study answers a narrower question than “Did teams get worse?”

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Does AI-generated code create more work for code reviewers?

It can, in principle, if AI increases the volume of changes faster than teams can assess them, or if reviewers need extra effort to verify generated code. But the cited studies do not directly measure the net effect on review hours, queue delays, or the amount of rework caused by AI-assisted submissions.

Approval likelihood and readability ratings are not substitutes for reviewer workload. A submission may be easier to approve while still requiring careful inspection; alternatively, more submissions may increase total review demand even if each individual review becomes faster. To answer the workload question, a team needs direct measures such as review time per accepted change and time waiting for a first review, tracked alongside the number and size of changes.

AI can also be applied to review practices themselves. Google Research’s AutoCommenter learns and enforces coding-language best practices, and its industrial evaluation reports a positive workflow impact. The public abstract does not give a numerical estimate for reviewer hours or defect outcomes, so it cannot establish whether automated feedback replaces, reduces, or simply changes human review.

Are developers spending more time reviewing code written by AI?

The available findings do not establish a general increase in time spent reviewing AI-written code. The METR trial measured time to complete assigned coding tasks, not time spent reviewing AI-generated changes. GitHub’s controlled study included blind reviewers, but its reported results do not provide an organization-wide accounting of review hours. Microsoft Research studied changes in developer practices and perceptions, rather than total review time across a software portfolio.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The question also depends on what counts as review. Time spent reading a diff, checking tests, revising a generated suggestion, addressing comments, and waiting in a review queue are different activities. A team that records only formal review duration could miss work shifted into author-side verification or rework.

Does AI coding make code quality worse?

There is no basis in these studies for saying that AI coding generally makes code quality worse—or that it makes quality better in production. GitHub’s results are evidence that, in one defined Python API task, its Copilot group performed better on specified tests and reviewer assessments. Those measures are useful within that exercise, but they do not tell us how code behaves after deployment, how maintainable it remains over time, or how often it contributes to incidents.

Quality has several dimensions that can move independently. A change can pass tests yet be difficult to maintain; read clearly yet introduce a security or reliability problem; or be accepted by a reviewer while increasing future maintenance burden. Evaluating quality therefore requires production outcomes and follow-up over time, not a single proxy such as task throughput, test passage, or approval.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How do you measure whether AI makes software teams more productive?

Measure delivery and consequences together, using a credible comparison over a defined period. Comparing AI users with non-users without accounting for task assignment, experience, project familiarity, or adoption choices can confound the result. Where feasible, use a randomized rollout or a carefully matched comparison, and report uncertainty as well as averages.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Define the comparison. Specify which tools and versions are in scope, which teams and tasks are eligible, and what counts as an accepted and shipped change.
  2. Track review effort and flow. Record reviewer hours per accepted change, time to first review, review-queue delay, comment severity, and rework cycles.
  3. Track production outcomes. Measure escaped defects, rollbacks, and incident severity per shipped change, as well as change failure rate.
  4. Include volume and change size. Report throughput alongside the size and complexity of changes. More merged changes alone do not establish more value if each change carries higher failure or maintenance costs.
  5. Measure lasting ownership. Assess maintenance burden and whether developers can explain, debug, and safely modify the code they own.
  6. Break down results. Separate findings by developer experience, familiarity with the codebase, task type, and AI-tool use so that an average does not conceal groups with different outcomes.

Set the measurement window to fit the outcome. Review time can be observed soon after a change, while escaped defects and maintenance burden may require longer follow-up. Report the full set of outcomes rather than combining unlike measures into a single productivity score.

What can we conclude about developers becoming reviewers?

The title’s “every developer” framing is stronger than the evidence supports. The studies document increased task completion in some field experiments, a slowdown in one small trial of experienced developers working in familiar repositories, quality results in a bounded vendor study, and shifts in developer perceptions. They do not quantify an economy-wide change in the balance between writing and reviewing code.

It is reasonable to investigate whether AI changes the amount or nature of review work. The unresolved question is not whether researchers have measured anything: they have measured task outcomes, bounded quality measures, approval, perceptions, and workflow effects. What remains unestablished is whether AI-assisted code production raises total review load or worsens long-term outcomes across software teams.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.