The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →AI coding assistants can make some software tasks faster, but the evidence does not support one productivity percentage for software engineering as a whole. In a controlled exercise, developers using GitHub Copilot finished a small JavaScript task sooner. In a randomized study of experienced contributors working on issues in repositories they knew well, developers allowed to use early-2025 AI tools took longer on average. The difference is a reminder that task, codebase, tool, and measurement all matter.
What the studies found
These studies measure different outcomes in different settings, so their results should be read side by side—not averaged into a universal estimate.
| Study and setting | What was measured | Result and scope |
|---|---|---|
| GitHub, 2022: a timed JavaScript exercise | 95 professional developers were randomly assigned to use Copilot or not while writing a JavaScript HTTP server. | The Copilot group averaged 1 hour 11 minutes, versus 2 hours 41 minutes without Copilot; task completion was 78% versus 70%. GitHub reported a 55% faster completion time and a 95% confidence interval of 21% to 89% for the speed gain. This is evidence about that bounded exercise, not a general estimate for software delivery. |
| METR, July 2025: issues in familiar open-source repositories | Sixteen experienced contributors submitted 246 real bugs, features, and refactors. Issues were randomly assigned to AI-allowed or AI-disallowed conditions. Tasks averaged about two hours; participants recorded screens and self-reported implementation time. In the AI condition, tool choice was open, with use primarily involving Cursor Pro and Claude 3.5 or 3.7 Sonnet. | Issues took 19% longer on average when AI was allowed. Before the study, participants expected a 24% speedup; afterward, they still believed they had been sped up by 20%. METR describes this as a snapshot of early-2025 tools and this particular setting, not proof that AI slows most developers. |
| UK Government Digital Service, November 2024–February 2025: public-sector trial | 2,500 licenses were made available across central government organizations and 1,900 assigned. The analysis included 424 survey responses from users in 31 departments; 73% reported at least five years of coding experience. GDS combined survey responses with usage telemetry. | 58% of respondents said they would not want to return to pre-assistant working conditions, and average satisfaction was 6.6 out of 10. Telemetry showed a 15.8% average acceptance rate for suggested code lines; 39% of respondents reported committing suggested code. These are sentiment, usage, and acceptance measures—not a randomized estimate of delivered output or end-to-end time saved. |
| GitHub, 2024, article updated February 2025: web-server API task | Developers with at least five years of experience were randomly assigned Copilot access or no AI; 202 valid submissions were analyzed. The task was assessed with ten unit tests and blind expert review. | GitHub reported that Copilot users were 53.2% more likely to pass all ten tests. This is a relative likelihood reported for the study task, not a 53.2 percentage-point increase. The company also reported favorable differences in functionality and expert-rated readability, reliability, maintainability, conciseness, and approval likelihood. |
| Microsoft Research, June 2025: three workplace trials | The publication describes randomized controlled trials at Microsoft, Accenture, and an anonymous Fortune 100 company. Random subsets of developers received access to an assistant with intelligent code completions. | The publication page establishes the trial settings and design but does not provide an outcome estimate in the material summarized here. It therefore cannot supply a numerical productivity result for comparison. |
Why the results differ
“Productivity” can mean finishing a narrowly defined task sooner, completing a task at all, producing code that passes tests, or improving delivery across a team. Those outcomes are related, but they are not interchangeable. A tool may help with a short, self-contained implementation and still add overhead when a developer must understand an established codebase, verify suggestions, and satisfy surrounding project conventions.
The METR study is particularly useful for illustrating that distinction: its participants were experienced contributors working in repositories they already knew, on real issues rather than isolated exercises. Its authors discuss the limits of generalizing from that setting, including the difference between realistic repository work and algorithmically scored benchmarks. The result describes early-2025 tools and does not establish how newer tools perform in 2026 or in other workflows.
#1 Best Overall
The positive task-speed and code-quality results also have boundaries. GitHub conducted those experiments and studied specific web-server tasks under controlled conditions. They show what happened in those tasks under those methods; they do not by themselves demonstrate better quality or faster delivery in production systems.
How to judge productivity claims
Before applying a result to a team, check whether the study resembles the work that team actually does. In particular, compare:
Rank #2
- Task type and complexity: a timed exercise, routine change, and multi-step issue in an existing service are different workloads.
- Codebase context: unfamiliar starter code is not the same as a mature repository with implicit conventions, tests, and documentation requirements.
- Participants: experience, repository familiarity, and prior use of an assistant can affect results.
- Tool and date: record the assistant, model, interaction mode, and when the test took place; findings about early-2025 tools are not timeless.
- Definition of success: elapsed time, completion rate, test results, expert review, suggestion acceptance, satisfaction, and organizational throughput answer different questions.
- Study design: distinguish randomized comparisons from workplace rollouts, telemetry, and self-reports, and note who sponsored or conducted the study.
What a team should measure in its own workflow
External studies provide useful signals, but a team deciding whether an assistant improves its delivery should evaluate it against its own tasks and definition of done. A practical comparison should include a representative mix of work, rather than only coding prompts that are easy to isolate.
- Choose representative tasks. Include the routine changes, bug fixes, tests, and repository-level work the team actually handles. Record enough context to compare task difficulty and developer familiarity.
- Compare similar conditions. Where feasible, use randomized or otherwise balanced assignments, and document which assistant and model versions are available. Avoid comparing one group’s familiar codebase work with another group’s unfamiliar exercises.
- Measure completed work, not typing alone. Track elapsed time through review and acceptance, completion, test outcomes, and defects or rework. Include tool usage and code acceptance as diagnostic measures, not as substitutes for delivery outcomes.
- Review quality independently. Apply the same tests and review criteria to assisted and unassisted work. A faster first draft is not a productivity gain if verification, rework, or downstream failures erase the time saved.
- Report uncertainty and scope. State how many tasks and developers were observed, what work was included, how success was defined, and whether results reflect an experiment or self-reported experience.
What the evidence supports
The defensible conclusion is conditional: AI coding assistants can improve speed or measured task quality in some controlled tasks, while they can also coincide with slower completion in complex, familiar-repository work. Positive user sentiment and code-suggestion acceptance show that developers find or use assistance; neither alone proves that the organization ships more reliable software faster. Teams should treat productivity as an outcome to measure in their own workflow, not an automatic property of adopting an assistant.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Quick Recap
Best Value
Rank #4
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

