Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

AI coding tools can help developers complete more work in some settings and slow them down in others. The apparent contradiction makes sense once productivity is measured as useful, reviewable software delivery—not as lines of code or the speed of generating a first draft.

Why more code is not the same as more productivity

Code volume is an activity measure. It says how much code was written, not whether the change solved the right problem, passed review, integrated cleanly, or remained maintainable. A tool can make code generation faster while leaving the harder parts of a task unchanged—or creating extra work elsewhere.

Productivity also has human and team dimensions. GitHub’s Copilot research uses the SPACE framework, which considers satisfaction and well-being, performance, activity, communication and collaboration, and efficiency and flow. The study examines only a subset of those dimensions, illustrating why a single output measure cannot stand in for the whole experience.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

That distinction matters when asking whether AI coding actually makes developers more productive. Faster typing, more generated code, perceived speed, completed tasks, and accepted changes are related but different outcomes. A result about one should not be presented as proof of another.

What the studies measured—and why their results differ

These studies do not form a direct contest. They used different developers, tasks, tools, settings, and definitions of success.

Study Participants and work Reported outcome What the result can establish
Microsoft Research, June 2025 Three randomized field experiments at Microsoft, Accenture, and an anonymous Fortune 100 company; 4,867 developers combined. 26.08% increase in completed tasks, with a standard error of 10.3%. In these participating workplaces and experiments, AI use was associated with an increase in the measured task-completion outcome. The pooled result is not a universal forecast for every team or task.
METR, July 2025 A randomized trial involving 16 experienced open-source developers and 246 issues in projects they knew; participants could use early-2025 AI tools. Issue completion took 19% longer in the trial. This is evidence about experienced developers doing realistic work in familiar, mature repositories—not a finding about all developers or all coding tasks.
GitHub, 2022; updated May 2024 A controlled JavaScript HTTP-server exercise with 95 professional developers using the Copilot version and setup in that experiment. GitHub reported 55% faster task completion. The result applies to that specific exercise and experimental context; it does not measure every stage of software delivery.

Microsoft Research also reported that less experienced developers had higher adoption and greater productivity gains in its experiments. That finding suggests experience may shape results, but it does not mean every less experienced developer will benefit more in every environment.

METR’s authors summarized their trial this way: “When developers are allowed to use AI tools, they take 19% longer to complete issues—a significant slowdown that goes against developer beliefs and expert forecasts.” The statement describes their trial, not a population-wide effect. Participants had expected a 24% speedup and, after the study, still believed they had been sped up by 20%, despite the measured slowdown. That gap is a reminder that subjective impressions and observed completion times can diverge.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why a coding benchmark can disagree with work in a real repository

A short, well-bounded exercise and an issue in a mature codebase place different demands on a developer. In the METR trial, work was intended to satisfy a human reviewer, including expectations around style, tests, and documentation. Many benchmarks instead score a result against test cases. Passing tests can be important, but it does not necessarily show that a change is easy to review, fits local conventions, or addresses implicit requirements.

That difference offers a plausible explanation for why code generation can look fast while end-to-end issue completion does not. Context loading, review, rework, specification ambiguity, or integration could consume time saved elsewhere. These are possible mechanisms, not established causes of the slowdown in the cited trial.

Familiarity with a codebase, the type of task, and the particular tools available can also change what help is useful. A tool that performs well when creating a small, isolated example may behave differently when a developer must navigate established architecture and satisfy requirements that are only partly written down.

What the evidence says about AI’s effect on teams

The results support a conditional answer, not a universal claim that AI always accelerates or always slows software development. Microsoft Research’s workplace experiments, METR’s open-source trial, and GitHub’s controlled exercise answer different questions. Their outcomes cannot be combined into one overall productivity percentage without a common measure and comparable conditions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Organizational context matters too. DORA’s 2025 report, published by Google Research, draws on more than 100 hours of qualitative research and responses from nearly 5,000 technology professionals worldwide. It describes AI as “an amplifier,” magnifying strengths in high-performing organizations and dysfunctions in struggling ones. That framing directs attention beyond the coding assistant: team practices and delivery conditions influence whether generated output becomes useful software.

GitHub’s research also includes qualitative testimony from a participant identified as a “Senior Software Engineer”: “(With Copilot) I have to think less, and when I have to think it’s the fun stuff. It sets off a little spark that makes coding more fun and more efficient.” This captures a reported experience, not a measured productivity result. Satisfaction and flow can matter to developers without proving that a team ships more accepted work.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to evaluate AI productivity in your own team

The practical question is not whether AI produces more code, but whether it helps your team deliver better changes with an acceptable amount of effort. The following is an evaluation approach derived from the distinctions in these studies, not a universal published standard.

  1. Define the outcome before trying a tool. Choose an outcome that reflects completed work in your process, such as accepted changes or issues completed to the team’s normal review standard. Track quality and reviewability alongside speed rather than treating code volume as the goal.
  2. Include the work after generation. Account for review, corrections, tests, documentation, and integration. A faster first draft is not an end-to-end gain if the added work erases the time saved.
  3. Separate unlike tasks. Compare similar work with similar work: for example, isolated exercises with isolated exercises, or repository issues with issues of comparable scope. Record the task type and codebase context so a result in one category does not conceal a different result in another.
  4. Segment by developer experience and tool use. Examine whether adoption and outcomes vary with experience, task, and the way the tool is used. An average can hide important differences across contributors or work types.
  5. Compare against a credible baseline over enough work. Use a comparison period or group that reflects normal work, and allow for variation and learning. A single task or a few enthusiastic users cannot show whether an apparent gain persists.
  6. Keep multiple signals visible. Consider delivery outcomes alongside developer experience, collaboration, and flow. If one signal improves while another worsens, report the trade-off rather than compressing it into a single productivity score.

How current is the METR slowdown result?

The 19% longer completion-time finding is from METR’s July 2025 trial using early-2025 AI tools. METR’s study page notes additional data on later-2025 tools published in February 2026. The July result should therefore be read as a dated study of its participants, tasks, and tools—not as a measurement of every current assistant or a forecast of what the same developers would experience today.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.