iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
No general 10x productivity gain is established by the available studies. They report different results for different kinds of work: a large speedup on one timed coding exercise, a more modest increase in completed tasks in workplace experiments, and slower completion in a trial of experienced developers working in familiar repositories. Those findings are not interchangeable—and none demonstrates a tenfold increase in developers’ overall, long-term output.
Why “productivity” needs a definition
A developer can finish one coding task faster without delivering more working software overall. The time saved might be offset by reviewing generated code, correcting mistakes, integrating changes, or maintaining them later. Conversely, a tool might help someone finish more tasks or make repetitive work feel easier without reducing the time for each task.
The studies below measure different outcomes: time to finish a bounded task, the number of workplace tasks completed, and developers’ own reports of focus or effort. They do not provide a common measure of quality-adjusted software delivery over the long term. Comparing their percentages as if they measured the same thing would give a misleading answer.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteWhat the studies found
| Study and setting | Participants and tools | Reported result | What the measure represents |
|---|---|---|---|
| GitHub’s controlled task experiment, described in a 2022 post updated in 2024; also reported by Microsoft Research in 2023 | 95 professional developers were randomly assigned Copilot access or no access. They implemented a JavaScript HTTP server, with correctness and completeness assessed by a test suite. | GitHub reported average completion times of 1 hour 11 minutes with Copilot and 2 hours 41 minutes without it, describing the result as 55% faster. Its reported 95% confidence interval was 21% to 89%, with P=.0017. Microsoft Research reported 55.8% faster completion for the same bounded task; the difference is a reporting or rounding variation, not a second experiment. | Time to complete one test-scored coding exercise—not speed across a developer’s whole job. |
| GitHub’s Technical Preview survey, described in the same 2022 post updated in 2024 | More than 2,000 developers who had signed up for the preview responded. | Depending on the statement, 60–75% agreed with positive claims about fulfillment, frustration, or focus; 73% said Copilot helped them stay in flow, and 87% said it preserved mental effort on repetitive tasks. | Self-reported experience among preview participants, not measured increases in completed work. |
| Three randomized workplace field experiments, summarized by Microsoft Research and published online in Management Science on February 27, 2026 | Pooled data from 4,867 developers at Microsoft, Accenture, and an anonymous Fortune 100 company; participants had access to an AI coding assistant. | The pooled estimate was a 26.08% increase in completed tasks, with a standard error of 10.3%. The researchers describe the individual experiments as noisy and their results as varying. Less experienced developers had higher adoption and larger productivity gains. | Completed task counts in those workplace experiments—not a universal individual speedup or a direct measure of long-term software value. |
| METR’s randomized trial, conducted from February to June 2025 | 16 experienced open-source developers completed 246 tasks in mature projects where they had an average of five years of prior experience. When AI was allowed, participants primarily used Cursor Pro and Claude 3.5/3.7 Sonnet. | AI access increased task completion time by 19%. Before starting, participants forecast a 24% time reduction; afterward, they estimated a 20% reduction despite the measured slowdown. | Task completion time in familiar repositories for this specific group and tool period. |
| METR’s later experiment, described in a February 2026 update | The experiment began in August 2025 and included 57 developers, 143 repositories, and more than 800 tasks. | METR said selection effects and unreliable time measurements for some participants using multiple agents made the experiment an unreliable signal of the current productivity effect. It reported raw estimates, including a speedup estimate for some returning developers, but said the design problems made the data weak evidence for the size of any increase. | A methodologically limited follow-up, not a settled estimate that resolves the earlier trial. |
Why the results differ
A self-contained exercise is not a normal development queue
The Copilot experiment asked developers to implement a defined feature as quickly as possible, with an automated test suite to assess the result. That makes it useful for answering whether assistance can help with a bounded coding task under timed conditions. It does not capture the full sequence of work on a production codebase: understanding requirements, navigating dependencies, reviewing changes, coordinating with teammates, and living with maintenance consequences.
#1 Best Overall
Workplace output depends on the tasks and the team
The field experiments measured completed tasks in three organizations, a closer fit to workplace activity than a standalone exercise. But the pooled estimate combines distinct experiments, and the researchers report that the results were noisy and varied. Task counts also do not, by themselves, establish that each task had equal difficulty, quality, or downstream value. The finding is evidence of gains in those settings, not a promise that another team will see the same effect.
Familiarity can change the value of assistance
METR’s early-2025 trial focused on experienced open-source developers working in mature projects they already knew well. Their repository knowledge may make some tasks faster without AI, while AI suggestions can still require evaluation and integration. The measured slowdown is important counterevidence to blanket claims, but its small, specific sample does not show that AI slows every developer or every kind of work.
Rank #2
Later data did not settle the current effect size
In its 2026 update, METR explained that some developers were increasingly unwilling to participate if they could not use AI, creating a risk of selection bias. It also said that using multiple agents made some participants’ time difficult to measure reliably. The raw estimates from that follow-up therefore should not be treated as a clean resolution of the earlier result or as a reliable current productivity multiplier.
Recommended Free Tools
How to judge a “10x” claim
Before accepting a productivity claim, ask what was measured and what work it covers. A useful comparison should make clear:
- The outcome: task time, number of completed tasks, correctness, quality, or self-reported experience.
- The task mix: a small, self-contained coding exercise, ordinary team work, or maintenance in a familiar codebase.
- The participants: their experience level and how well they know the repository.
- The tool period: which assistant and model versions were available when the study took place.
- The study design: randomized comparison, field experiment, or survey—and the uncertainty or measurement limitations the authors report.
These distinctions matter because a narrow task-speed result cannot establish that a developer delivers ten times as much reliable software over months of work. A survey about focus cannot be converted into a percentage increase in output, and a result from one group or tool period should not be silently generalized to another.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.What the evidence supports
The most defensible conclusion is conditional: AI coding assistance can improve measured outcomes in some settings, but the size and direction of the effect vary with the work, the developers, the tools, and the metric. The evidence reviewed here does not establish a general 10x increase in developer productivity, nor does it show that no one could see a dramatic gain on a carefully chosen task. It is not a head-to-head comparison of the same developers, tasks, and tools, so the results should be read as evidence about different settings rather than averaged into one universal multiplier.
Quick Recap
Best Value
Rank #4
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

