Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

AI coding assistants can speed up some software tasks, but that does not by itself prove they reduce the cost of maintaining software. A real saving depends on the work needed to review, test, correct and later change the code—as well as on the team’s existing practices. Current evidence includes faster initial task completion and favorable results on a narrow code-quality test, but a controlled follow-up study found no significant difference in the effort or quality of later changes.

What counts as software maintenance cost?

Maintenance is the work of changing software to improve it, fix it or adapt it. ISO 25010 defines maintainability as “the degree of effectiveness and efficiency with which a product or system can be modified to improve it, correct it or adapt it to changes in environment, and in requirements.” Borg and colleagues quote this definition in their 2026 paper in Empirical Software Engineering.

For a team, the cost of a change is therefore more than the time spent writing code. It can include:

  • Understanding the existing system and locating the right place to make a change.
  • Reviewing the proposed change and correcting errors or unclear design choices.
  • Writing and running tests, investigating failures and fixing regressions.
  • Supporting the code in later work, when another developer must modify or diagnose it.

Technical debt is a useful way to describe one risk in this equation: a design or implementation choice that is expedient now but can make later changes more costly or even impossible. AI assistance does not automatically create technical debt; the concern is whether a particular output makes future work harder.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Where AI may save effort—and what that does not prove

An assistant may help draft routine code, suggest an implementation, or provide a starting point for tests or review. If the suggestion is correct and fits the codebase, it can reduce the time spent on the immediate task. But faster completion of that task is only one part of the maintenance bill. Review, rework, defect correction and later change effort still matter.

A UK government trial reported by IT Pro in 2025 illustrates the distinction. More than 1,000 workers across 50 UK government departments tested tools from Microsoft, GitHub and Google between November 2024 and February 2025. The report said developers saved around one hour per day, equivalent to around 28 working days per year, and that 15% of AI-generated code was used without edits. These are reported trial figures, not a controlled estimate of long-term maintenance savings; the editing figure also shows why generated code should not be treated as finished work.

What studies say about code that must be maintained later

Evidence What was studied Reported result What it does not establish
Borg et al., Empirical Software Engineering, 2026 A two-phase study using a Java web application task. The controlled study included 151 participants, 95% of them professional developers. In Phase 1, developers added a feature with or without AI assistance; in Phase 2, new participants evolved the resulting solutions without AI assistance. Phase 1 observational results showed a 30.7% median reduction in completion time for the initial feature task. In Phase 2, the authors found no significant difference in subsequent developers’ completion time or code quality. Bayesian analysis indicated that any speed or quality effects were small and uncertain. A universal effect on maintenance costs, or proof that every AI-generated solution is equally maintainable. The authors also note that autonomous coding agents were not represented in the empirical results from tools available in late 2024.
GitHub, company-authored study published in 2024 and updated in 2025 A randomized study with 202 experienced developers in the final valid sample, completing one API-endpoint task. The study assessed unit tests and expert review. GitHub reported statistically significant improvements on several measured code-quality dimensions, including a 2.47% improvement in maintainability ratings in its task-specific comparison. Lower maintenance bills over months or years. One task and a small final sample cannot establish long-term costs, and the study was authored by the tool vendor.
DORA / Google, 2025 A report drawing on more than 100 hours of qualitative data and responses from nearly 5,000 technology professionals. DORA described AI as an amplifier of organizational strengths and weaknesses: “It magnifies the strengths of high-performing organizations and the dysfunctions of struggling ones.” A randomized estimate of how much an individual team’s maintenance costs will change because it adopts an AI assistant.

Taken together, these findings do not support a single percentage reduction in software maintenance costs. They measure different outcomes in different settings: initial task speed, ratings on a specific coding task, later developer effort, and organization-wide experience. In particular, the Borg study’s follow-up is evidence about its bounded Java task and tools available at the time—not a verdict on all codebases or today’s autonomous agents.

Why the team’s working system matters

AI output enters an existing delivery process. If a team has useful tests, clear review responsibilities and short feedback loops, it has ways to identify and correct unsuitable code. If tests are weak or reviewers lack time and context, faster production of code can leave more uncertainty to resolve later. This is the practical significance of DORA’s 2025 “amplifier” framing: tool choice alone does not determine the outcome.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Test coverage: Tests help reveal whether a change breaks expected behavior. They do not guarantee good design, so passing tests should not replace review.
  • Review quality: Reviewers need to assess correctness, fit with the codebase and ease of future modification—not just whether the code looks plausible.
  • Feedback and rework: Track time spent correcting generated suggestions and resolving failures. An apparent speed gain can be offset if those costs are high.
  • Tool generation and task type: Results from a single feature or endpoint task should not automatically be generalized to unfamiliar legacy code, complex changes or autonomous agents.

How to measure whether AI reduces your maintenance costs

Compare similar work with and without AI assistance, and account for the full change process rather than relying on a single productivity figure.

  1. Choose comparable tasks. Match work by type and difficulty as closely as practical, and record whether the task involved new code, an existing component or a legacy system.
  2. Record immediate completion time. Measure from the start of the task to a review-ready change. Keep this separate from time spent typing or accepting suggestions.
  3. Count review and rework. Record reviewer effort, corrections, test failures and time spent fixing problems before the change is accepted.
  4. Check quality outcomes. Track functional correctness, regression-test results and defects found after release. Do not treat a code-quality rating as a direct measure of dollars saved.
  5. Measure later changes. When comparable follow-up work occurs, record how long another developer needs to understand and modify the code, and whether the change introduces new problems.
  6. Compare the total effort. Evaluate immediate work, review, correction, defect response and later changes together. Keep the measurement period and task mix visible so a short-term result is not mistaken for a lasting one.

Maintainability measures can add context to time and defect data. For example, Borg and colleagues describe CodeScene’s CodeHealth metric as measuring code smells. Such a signal can help identify areas for inspection, but it should not be presented as a direct cost measure unless the team has established how it relates to its own maintenance work.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Using AI with legacy code

Legacy maintenance often begins with understanding dependencies and creating enough feedback to make a safe change. AI may help explain unfamiliar code or suggest a patch, but those suggestions do not remove the need to verify behavior and preserve the ability to change the system safely.

Michael Feathers’s Working Effectively with Legacy Code is a relevant, non-AI-specific reference for this work. Pearson’s paperback listing covers feedback, tests, dependencies, safe changes and refactoring. The book was published in 2004, so it should be read for its legacy-code techniques, not as guidance about generative AI tools.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When AI is worth adopting for maintenance work

Adopt an assistant for a maintenance workflow when measurements show that it reduces net effort without weakening safeguards. A faster first draft is a promising signal, not the final decision: the result should also hold up under review, testing and subsequent changes. If the team cannot yet measure those later costs, treat claims of long-term savings as unproven rather than assuming that task-speed gains will compound.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.