Free tools Windows power users keep installed
One-click scans. No signup required.
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
AI coding tools can help developers complete more work or produce code that fares better on a particular test. They can also add time when a task demands deep familiarity with a mature codebase. The useful management question is not whether AI makes every developer faster, but whether a team can turn generated code into work that is correct, reviewable, and maintainable.
Does AI actually make software developers more productive?
Sometimes, but the answer depends on what “productive” means and what developers are doing. Studies have measured elapsed time on a bounded coding task, completed tasks in company settings, performance on unit tests, reviewer ratings, and developers’ own impressions. Those are different outcomes, not interchangeable estimates of one universal productivity gain.
DORA’s 2025 report offers a useful organizational frame: it describes AI as an amplifier that magnifies strengths in high-performing organizations and dysfunctions in struggling ones. The report combines more than 100 hours of qualitative data with survey responses from nearly 5,000 technology professionals worldwide. That is DORA’s synthesis of its research, not a controlled estimate of how much AI speeds up a weak team—or proof that any particular engineering practice causes better AI results. DORA 2025 report, Google Research
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
What the studies measured
| Study and setting | Reported result | What that result can tell you |
|---|---|---|
| Microsoft Research, three randomized field experiments at Microsoft, Accenture, and an anonymous Fortune 100 company; 4,867 developers, 2025 | The authors estimated a 26.08% increase in completed tasks for developers given an AI coding assistant (standard error: 10.3%). They describe the individual experiments as noisy and report larger adoption and productivity gains among less-experienced developers. | A combined estimate of completed work in those participating companies—not a universal time saving or guarantee for another team. Microsoft Research study |
| METR, randomized study of 16 experienced contributors to large open-source repositories; 246 issues, July 2025 | Participants took 19% longer to complete issues when allowed to use AI tools. | A result for experienced maintainers working in this study’s setting with early-2025 tools—not evidence that AI slows most developers or every kind of task. METR study |
| Microsoft Research, controlled experiment in which developers built a JavaScript HTTP server, 2023 | The group using GitHub Copilot completed the task 55.8% faster than the control group. | A result for a tightly scoped implementation task and an earlier tool generation, not a forecast of team-wide productivity. Microsoft Research study |
The contrast is real, but it is not a clean contradiction. The workplace trials counted completed tasks across company settings; METR timed work on issues in repositories where contributors were experienced; the earlier Copilot experiment measured the time to implement a defined server. Task context, tool generation, developer experience, and evaluation method all differ.
#1 Best Overall
Why do AI coding productivity studies disagree?
AI can draft code quickly without removing the work needed to understand a system, resolve ambiguous requirements, test a change, and make it acceptable to reviewers. A self-contained programming exercise can reward fast code production. A change in an established repository may depend on conventions and dependencies that are not obvious from the issue description. The same tool can therefore help with one part of the work while adding little—or even adding friction—to another.
Compare the conditions, not just the headline percentages
- Task and repository: A short implementation exercise is not equivalent to debugging, refactoring, or adding a feature to a mature codebase.
- Developer experience: The three-company Microsoft field-trial abstract reports larger gains among less-experienced developers. METR’s study focused on experienced maintainers.
- Tool and date: METR’s findings describe early-2025 tools, primarily Cursor Pro with Claude 3.5 or 3.7 Sonnet, with participants choosing the tools they used. Capabilities change, so the result should not be treated as a verdict on later systems.
- Outcome: Elapsed time, completed-task counts, unit-test results, code-review ratings, and perceived usefulness answer different questions.
- Quality threshold: A task that is complete only after it passes tests and meets review, style, or documentation expectations may take longer than one scored by an automated benchmark.
METR notes that realistic pull requests can require developers to satisfy human reviewers on style, testing, and documentation, unlike benchmark tasks scored algorithmically. Its participants expected a 24% speedup and, after taking longer in the experiment, still believed AI had sped them up by 20%. That gap is a reason to distinguish perceived speed from measured time, not to dismiss developers’ experience of usefulness. METR’s study and limitations
Rank #2
Does AI-generated code have lower quality?
The available evidence here does not support a blanket claim that AI-generated code is lower quality—or that it is universally better. A randomized GitHub study recruited 243 developers with at least five years of Python experience; 202 valid submissions were analyzed, with 104 developers using Copilot and 98 not using it. Participants implemented API endpoints for a fictional restaurant-review web server. The Copilot group had a 53.2% greater likelihood of passing all ten unit tests. In blind reviews, ratings were also higher for readability by 3.62%, reliability by 2.94%, maintainability by 2.47%, and conciseness by 4.16%; the group had a 5% higher likelihood of approval. These are results from that controlled exercise and its measures, not evidence about long-run production defects or maintenance costs. GitHub Research study, published 2024 and updated 2025
There is an important boundary to the study’s quality claim: GitHub’s rubric counted issues such as unclear identifiers, missing documentation, repeated code, and excessive branching as code errors, but did not count functional errors that prevented code from working. Passing the ten tests and receiving favorable ratings under that rubric are useful evidence about the exercise; neither guarantees that generated code is safe, correct in every case, or easy to maintain in a production system.
Can AI fix weak engineering practices?
AI can suggest code, tests, explanations, and changes. It cannot by itself establish that a team has clear requirements, reliable ways to detect regressions, a sound review process, or shared understanding of its architecture. If those foundations are weak, faster code production may simply move the bottleneck to verification, integration, or rework. That is a practical interpretation of DORA’s amplifier framing, not a demonstrated causal finding that weak practices always make AI outcomes worse.
The studies summarized here do not prove that any one practice—such as writing tests, conducting code review, or improving documentation—causes larger gains from AI. Teams should treat engineering discipline as a way to make changes observable and reviewable, not as a guaranteed productivity multiplier. A tool that produces more code is not automatically producing more value if a team cannot tell whether the code meets its needs.
Rank #4
How to judge AI’s value on your team
Use a small, defined evaluation that reflects the work your developers actually do. Keep the comparison fair: identify which tasks are eligible for AI use, what tools and versions are involved, and what counts as a finished change. Track more than how quickly a first draft appears.
- Choose representative work. Include tasks from the repositories and workflows where the team expects to use AI, rather than relying only on toy examples.
- Set completion criteria in advance. Define what “done” means, including relevant tests and review requirements, so the measure is not just code generated or time to first draft.
- Record the context. Note developer experience, task type, repository familiarity, tool and model generation, and whether AI was allowed or used.
- Measure separate outcomes. Compare time and completed work alongside test results, review feedback, rework, and developer experience. Do not merge them into one productivity score without a defensible method.
- Look for bottlenecks. If drafting gets quicker but review, debugging, or integration takes longer, the tool may have shifted work rather than reduced it.
- Reassess as tools and workflows change. Results from an older model or a different kind of task may no longer describe the team’s current use.
Workplace perceptions matter too, but they should be read as perceptions. In a separate Microsoft Research study at a large multinational software company, participants reported increased perceptions of usefulness and enjoyment with sustained use, while trust in AI-generated code did not increase. The study reports that 84% saw positive changes in daily work practices and 66% noted changes in how they felt about their work; these are participant reports, not measured output or code-quality gains. Microsoft Research workplace study
Quick Recap
Best Value
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

