iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
AI can produce code faster without making software delivery faster. The apparent gain can be consumed by prompting, review, rework, integration, testing, and maintenance—and whether it is depends on the task and the team’s workflow.
What “faster” means—and what it does not
Code-generation speed is a local measure: how quickly a tool produces a draft or a developer writes a change. Task completion includes understanding the request, checking the result, revising it, and getting it to work. Team delivery adds integration, testing, release, and coordination. Long-term maintainability concerns whether future developers can safely understand and change the software.
Those outcomes are related, but one does not prove another. A tool can make a first draft arrive sooner while increasing the work needed to verify or integrate it. Conversely, a faster workflow on one kind of task does not establish that every task, team, or codebase will benefit.
What the studies show
METR: a measured slowdown in one specific setting
In an early-2025 randomized trial, 16 experienced open-source developers completed 246 tasks in mature projects they already knew well. For the tasks in that study, allowing AI tools increased measured completion time by 19%. The result is a finding about that participant group, task set, and project context—not a universal estimate for developers or current AI coding work. METR’s study abstract reports the experiment.
#1 Best Overall
The mismatch between expectation and measurement is notable. Before the experiment, participants expected AI to cut completion time by 24%; afterward, they estimated a 20% reduction, even though recorded task time had increased by 19%. These are participants’ forecasts and retrospective estimates, not measured productivity gains.
METR’s 2026 update: the result is harder to interpret, not reversed
In a February 24, 2026 update, METR reported the earlier estimate as a 19% slowdown with a confidence interval from a 2% to a 39% slowdown. The organization also explained why later estimates were difficult to interpret: developers and tasks expected to benefit most were more likely to be excluded from the experiment, while concurrent use of agents complicated time measurement. METR cautioned that these factors could make observed results understate productivity uplift. That warning limits what can be inferred from later raw estimates; it is not conclusive proof that AI now speeds up work. METR’s update describes the design and measurement issues.
DORA: organizational conditions shape the outcome
DORA’s 2025 report characterizes AI as an amplifier of an organization’s existing strengths and weaknesses. That framing shifts the question from “Does the tool write code quickly?” to “Can the team turn more generated change into reliable, useful software?” Strong practices can help a team absorb and validate changes; weak review, testing, or integration practices can magnify the burden. DORA’s 2025 report discusses AI in the context of the wider delivery system.
Free tools Windows power users keep installed
One-click scans. No signup required.
Microsoft Research: positive perceptions are not proof of faster delivery
A 2025 mixed-methods study at a large multinational software company found that sustained use of generative AI coding tools led to more positive views of usefulness and enjoyment, while views of generated-code trustworthiness remained unchanged. In that study, 84% of participants reported positive changes in their daily work practices. That is a participant-reported perception result, not a measured productivity effect or evidence that code became more trustworthy. Microsoft Research’s study page describes the findings.
Rank #3
Where the apparent time saving can go
The following are useful places to look when a team sees faster generation but little improvement in delivery. They are diagnostic questions, not a validated scorecard with universal thresholds.
- Task familiarity and codebase maturity: Is the developer working in a project they understand, or does each change require learning unfamiliar conventions and dependencies?
- Prompting, review, and rework: How much time is spent specifying context, checking generated code, correcting mistakes, and revising a change before it is acceptable?
- Testing and documentation: Are tests and documentation keeping pace with the volume of changes, or is verification becoming a bottleneck?
- Integration and release: Do changes merge and ship smoothly, or do they create more conflicts, review queues, or coordination work?
- Team capacity: Can the team’s existing processes absorb more proposed changes without weakening review or reliability?
These questions help distinguish a local gain from an end-to-end one. If code arrives faster but waits longer for review or requires more correction, generation time may improve while delivery time does not.
Rank #4
How to judge whether AI is saving your team time
- Choose a meaningful unit of work. Compare tasks that resemble the work your team actually does, rather than counting generated lines or measuring only the time to produce a first draft.
- Track the whole path to completion. Include the time spent understanding the task, prompting, editing, reviewing, testing, integrating, and resolving follow-up issues.
- Separate kinds of outcomes. Record task completion and delivery measures separately from developer perceptions, code correctness, and maintenance outcomes; none is a substitute for the others.
- Compare like with like. Account for task familiarity, project maturity, and workflow differences before attributing a change to AI use.
- Revisit the result as tools and usage change. Newer agentic workflows and changing user selection can make older estimates poor guides to a team’s current experience.
What is not established about maintenance cost
The studies summarized here do not establish a universal long-term increase in maintenance cost or technical debt caused by AI-generated code. Faster code generation may raise a reasonable question about whether review and upkeep can keep pace, but the evidence described above does not quantify that effect. Treat maintenance as an outcome to monitor in your own software, not as a known percentage to apply to AI-assisted work.
Quick Recap
How to read the evidence without overgeneralizing
- METR’s 2025 experiment provides a measured result for experienced developers working in familiar, mature open-source projects; it does not predict every developer’s result.
- METR’s 2026 update highlights selection and timing problems that make later estimates difficult to interpret; it does not erase the earlier experiment or settle the productivity question.
- DORA’s systems framing explains why organizational practices matter, but it is not a numerical forecast for an individual team.
- Microsoft Research’s 84% result concerns reported changes in daily work practices, not measured speed or code quality.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

