Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

AI coding tools can make some development work slower, even when they generate code quickly. In METR’s 2025 randomized trial, experienced contributors took 19% longer on assigned tasks when they could use early-2025 AI tools. Other studies found productivity gains in different settings. The practical lesson is not that AI always helps or hurts: it is that task, codebase familiarity, verification effort, and the way a team measures work all matter.

Why can AI make coding slower?

Generating a plausible patch is only one part of finishing a software task. A developer may also need to explain the codebase to the tool, assess whether its assumptions fit the project, correct mistakes, integrate the change, and retain enough context to maintain it. These are plausible workflow costs to investigate—not a proven explanation for a specific share of METR’s measured slowdown.

The setting matters. METR studied experienced open-source developers working in mature repositories they already knew. For those tasks, success meant producing a change a human user would be satisfied could pass review, including the project’s expectations for style, testing, and documentation. That is different from a short, isolated exercise judged mainly by whether tests pass.

AI can therefore save typing while adding work elsewhere. If a change is deeply dependent on local conventions or has a high review burden, generated code may take longer to verify and integrate than writing the change directly. Whether that happens in your own workflow is something to measure, not assume.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the studies actually found

The findings are not directly contradictory: they examine different developers, tasks, tools, and definitions of productivity. A percentage for time per exercise cannot be compared as if it measured the same thing as completed tasks in workplace settings.

Study Setting and tool Reported result What the result does—and does not—show
METR, 2025 Randomized trial with 16 experienced developers and 246 tasks in mature open-source repositories they knew well; early-2025 tools, mainly Cursor Pro and Claude 3.5/3.7 Sonnet. Tasks took 19% longer with AI enabled. Participants expected a 24% time reduction before the trial and estimated a 20% reduction afterward. Evidence of a slowdown in this specific maintenance-task setting—not a forecast for every developer, task, or current tool.
Microsoft Research, 2025 Three randomized field experiments at Microsoft, Accenture, and an anonymous Fortune 100 company, covering 4,867 developers. The combined estimate was a 26.08% increase in completed tasks; individual experiments were noisy, and gains were greater among less experienced developers. Completed-task counts in workplace experiments are not the same outcome as time per METR maintenance task.
GitHub, 2022 Randomized exercise: 95 professional developers built a JavaScript HTTP server, with or without GitHub Copilot. The Copilot group averaged 1 hour 11 minutes versus 2 hours 41 minutes for the control group; GitHub reported 55% faster completion. A bounded exercise with a specific product and outcome; it does not establish the same effect for codebase maintenance. GitHub is the product vendor.
GitHub, 2024, updated 2025 Separate randomized code-quality study: 202 experienced developers submitted code for a web-server API exercise using Copilot or a control condition. GitHub reported a 53.2% greater likelihood of passing all 10 unit tests for the Copilot group, as well as small gains on several expert-rated quality dimensions. A study-specific exercise result, not a general estimate of production defect rates or maintainability.

“Productivity” can refer to elapsed task time, number of tasks completed, code quality, perceived effort, or end-to-end delivery time. Those measures can move in different directions. Do not average these study percentages into one general claim about how much AI increases or decreases productivity.

How to tell whether AI fits a coding task

Before handing work to an assistant, check the task and its verification costs. Assistance is more promising when it can produce a useful draft, explanation, repetitive transformation, or starting point that is cheap to check. A broad change in a mature, unfamiliar-to-the-tool codebase deserves more caution: the tool may lack the context needed to follow local conventions, and validating the result may be harder than making the change yourself.

  • Task shape: Is this a bounded exercise, a new feature, a bug fix, or a maintenance change embedded in established behavior?
  • Context: Can you identify the relevant files, constraints, and existing patterns clearly?
  • Verifiability: Can you check correctness with tests, a diff review, and project-specific expectations?
  • Definition of done: Does the work need only to pass tests, or also to meet review, style, documentation, integration, and maintenance requirements?
  • Tool and mode: Which product, model generation, and interaction mode are you using? Findings from early-2025 tools or a particular Copilot experiment do not automatically apply to a different tool or mode.

A workflow that keeps AI assistance reviewable

  1. Choose the task first. Identify a discrete part where assistance has a plausible benefit and where you can judge the answer. Treat highly contextual changes as an experiment rather than an automatic AI task.
  2. Give bounded context and a specific request. Name the relevant files, constraints, expected behavior, and tests. Ask for a small change that can be reviewed; avoid inviting a broad rewrite unless the scope genuinely calls for one.
  3. Keep verification in the task. Run relevant tests, inspect the full diff, compare assumptions with the codebase, and meet the same review and documentation bar you would use for unaided work. A generated patch is not finished work simply because it compiles or passes one test.
  4. Measure the whole job. For similar tasks, compare assisted and unassisted work while counting context-setting, correction, review, integration, and follow-up time—not just code generation or typing. Track quality and developer experience separately from elapsed time.
  5. Make it easy to switch back. If context costs rise or the output becomes harder to verify than the change itself, stop using the assistant for that task and continue directly. Review effects at team level as well as individual task speed.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Why team conditions matter

A tool does not operate independently of the engineering system around it. DORA’s 2025 report puts it this way: “AI’s primary role is as an amplifier, magnifying an organization’s existing strengths and weaknesses.” Clear ownership, reliable tests, manageable review queues, and shared conventions can make useful output easier to validate; weak foundations can make errors and coordination costs harder to contain. The report’s framing is organizational, not a guarantee that any particular tool or process will improve delivery.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a team, the useful question is not simply whether developers use AI. It is whether the full path from request to accepted, maintainable change improves without shifting hidden work into reviews, integration, or follow-up. METR’s preprint is a narrow snapshot of early-2025 tools and a specific group of contributors; the Microsoft and GitHub findings describe other settings, not a universal counter-result.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.