Sometimes, AI can help developers produce more work while leaving them less able to explain or solve what they just built. But the evidence does not prove that AI broadly makes engineers worse or employees better: studies measure different tasks and outcomes, and the learning evidence is short-term.
What do the studies actually show?
The findings point in different directions because the experiments asked different questions. A quiz after learning a new library measures something different from task throughput at work or the time needed to resolve issues in a mature codebase.
| Study and setting | What was measured | Reported result |
|---|---|---|
| Anthropic, 2026: 52 mostly junior developers learning the unfamiliar Trio Python library | Immediate quiz performance and task completion time | Average quiz scores were 50% with AI assistance and 67% with hand-coding, a 17-percentage-point gap. The AI group finished about two minutes faster on average, but that difference was not statistically significant. |
| Microsoft Research, June 2025: randomized field experiments at Microsoft, Accenture, and an anonymous Fortune 100 company; 4,867 developers combined | Completed tasks | AI-assistant access was associated with 26.08% more completed tasks across the experiments; individual results were noisy. |
| Bank for International Settlements, 2024: CodeFuse field experiment at Ant Group | Lines of code | The treatment group produced 55% more lines of code; statistically significant gains were primarily among junior employees. |
| METR, July 2025: randomized trial with 16 experienced open-source developers working on 246 issues in repositories they knew well | Issue completion time | Developers took 19% longer with the early-2025 AI tools tested in this setting. |
These percentages cannot be combined into one general “AI productivity” figure. More completed tasks, more lines of code, faster issue completion, and a higher quiz score are not interchangeable measures. None by itself establishes that code was better, that users learned more, or that an organization created more value.
Does AI make software engineers worse at coding?
Anthropic’s randomized study offers a reason to take the learning question seriously, but it does not show a permanent decline in engineering ability. Participants were familiar with Python but not Trio, an asynchronous Python library. They completed two feature tasks in an online coding platform and then took a quiz covering debugging, code reading, code writing, and concepts. The largest score gap was on debugging questions.
#1 Best Overall
The result concerns immediate mastery after a short learning exercise, not how these developers performed months later or across their jobs. Anthropic’s authors say the study does not resolve whether quiz performance predicts longer-term skill development. Its relatively small, mostly junior sample and sidebar-style assistant also differ from agentic coding products and many real workplace workflows.
That boundary matters: a developer can deliver a working feature with assistance without having built the same mental model as someone who reasoned through it unaided. Conversely, one short-term quiz cannot establish that using AI causes lasting deskilling. The study identifies a possible trade-off worth monitoring, not proof of irreversible damage.
Do AI coding assistants improve productivity?
Some workplace experiments report higher throughput, particularly for less experienced developers, but their measures and conditions matter. Microsoft Research’s three-company study counted completed tasks; the BIS summary counted lines of code. Neither figure is a quality-adjusted productivity measure, and more code is not automatically better code.
The counterexample from METR involved experienced developers addressing real issues in large open-source repositories they had contributed to for years. That work depends on navigating established code, constraints, and local conventions; it is not equivalent to a small, self-contained coding task. METR cautions that its result does not represent all software development. The study is also a snapshot of early-2025 tools, so it should not be generalized to later tools without evidence.
Recommended Free Tools
The fair conclusion is conditional: AI assistance has been associated with greater task output in some company experiments and slower issue completion in one demanding open-source setting. The results do not establish that either pattern applies to every developer, task, or tool.
Can AI help developers work faster while weakening their skills?
Yes, those outcomes can coexist: finishing a task and learning how to do it are separate goals. The Anthropic experiment makes that distinction visible, while the workplace productivity studies do not test whether their participants gained or lost skills. In particular, productivity gains among junior staff do not show that those workers were deskilled.
Rank #4
Anthropic’s analysis of screen recordings found that participants who asked conceptual questions or sought code explanations appeared among higher-scoring groups, while heavy delegation and AI-led debugging or verification appeared among lower-scoring groups. The authors explicitly caution that this qualitative analysis does not establish that those interaction patterns caused the score differences. Treat them as observations that suggest questions for training—not as proven rules for getting better results.
Does AI make employees feel better about work?
Workplace experience is another distinct outcome. Microsoft Research’s 2025 “Dear Diary” study at one large multinational software company combined surveys, an RCT, and a three-week diary study. It reported more positive perceptions of usefulness and enjoyment as employees introduced and sustained tool use, while perceptions of code trustworthiness did not change.
Best Value
In that study, 84% of participants reported positive changes in their daily work, and 66% reported some change in how they felt about work. These are participant reports from one company, not objective productivity measures or representative statistics for the software workforce. The diaries included enthusiasm as well as heightened pressure to keep up with new tools, so a more positive view of usefulness is not the same as an across-the-board wellbeing improvement.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How can teams use AI without outsourcing engineering judgment?
The studies do not validate a single learning-preserving workflow. A prudent approach is to keep developers responsible for understanding and verifying the work, especially when the task is intended to build new skills.
- Use assistance deliberately. Decide whether the immediate goal is delivery, learning, or both; a workflow optimized for completion may not require the same effort as one designed to teach.
- Ask for explanations as well as output. Have the assistant explain unfamiliar code or concepts, then check that explanation against the code and relevant documentation.
- Keep debugging and review in the loop. Trace failures, inspect changes, and run appropriate tests rather than treating generated code or AI verification as proof of correctness.
- For junior developers, pair delivery with practice. Teams can reserve time for unaided problem-solving, code review, and discussion of why a change works. This is sensible training practice, not an intervention whose impact these studies quantified.
- Measure the outcome that matters. Track quality, maintainability, learning, or task completion separately instead of treating code volume as a proxy for all of them.
Teams seeking structured learning can pair AI use with programming or software-engineering fundamentals; the cited studies do not endorse any particular course or book.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

