Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Successful software quality is not captured by one score. Teams need to see both how efficiently they deliver changes and how often those changes create disruption, then add product-quality signals such as defects found after release and what automated tests actually exercise. The seven measures below are a practical combination—not an official standard—and are most useful when defined for one service, interpreted in context, and followed over time.

Why use seven measures instead of one score?

DORA’s current software delivery framework defines five measures: three for throughput and two for instability. Escaped defects and automated test coverage are added here as complementary quality signals drawn from U.S. Department of Defense software metrics guidance. Together, they help answer two distinct questions: how does work move into production, and what happens to the product and its users?

DORA describes its delivery metrics as focusing on a team’s ability to deliver software safely, quickly, and efficiently. Its guide also frames them as leading indicators for organizational performance and employee well-being, and lagging indicators for software development and delivery practices. These are signals to prompt investigation—not a universal rating of quality.

The five software delivery measures

DORA groups its measures into throughput and instability. Read them together: higher deployment frequency or shorter lead time is not evidence of success if failures and unplanned work are also rising.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

1. Change lead time

Change lead time is the elapsed time from a change being committed to version control until it is deployed in production, as defined in DORA’s metrics guide. It helps reveal delays across the delivery system. Agree on which changes and production deployments count, and use the same start and end points consistently.

2. Deployment frequency

Deployment frequency measures how often a service is deployed to production over a period; teams may also examine the time between deployments. DORA’s definition and cautions make clear that frequency alone does not establish quality. A higher rate can coexist with instability, so pair it with change fail rate, deployment rework rate, and user-impacting defect information.

3. Failed deployment recovery time

This is the time needed to recover from a failed deployment that requires immediate intervention. DORA’s current label, failed deployment recovery time, is narrower than a general “mean time to recover”: it concerns a failed deployment, not every kind of service incident. Define when recovery starts and ends so the measure is comparable across the same service over time.

4. Change fail rate

Change fail rate is the proportion of deployments that require immediate intervention after deployment, such as a rollback or hotfix, according to DORA. It is a rate, so report both the numerator (interventions) and denominator (deployments) and state what qualifies as immediate intervention. A percentage without its counting rules can conceal important differences.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

5. Deployment rework rate

Deployment rework rate is the proportion of unplanned deployments made because of a production incident, per DORA’s definition. It captures a different form of instability from change fail rate: a deployment can lead to corrective unplanned work even when the original deployment was not counted as requiring immediate intervention. Keep the two measures distinct rather than merging them into a single failure figure.

Two complementary product-quality signals

6. Escaped defects

Escaped defects are defects discovered after release or outside the phase in which the team expected to catch them. The U.S. Department of Defense’s April 2023 software metrics guide includes escaped defects among software quality measures. Before tracking them, decide locally what counts as a defect, which release or testing boundary it escaped, how severity is recorded, and how long after release findings are attributed. Those choices affect the result; an unlabeled total can obscure whether user risk is improving.

7. Automated test coverage

Automated test coverage describes the portion of code or behavior exercised by automated tests, using a stated and consistently applied method. The DoD guide lists automated test coverage as a software metric, while DORA’s continuous-delivery guidance emphasizes that effective suites find real failures and only pass code that is releasable.

Coverage is evidence about test reach, not proof that tests detect meaningful faults. A high coverage figure can still come from tests that exercise code without checking important outcomes. Interpret it alongside escaped defects and the team’s evidence that its tests catch failures relevant to users.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to make the metrics useful

Define boundaries and denominators

Write down what counts as a change, production deployment, intervention, unplanned deployment, escaped defect, and covered code or behavior. For rates, retain the underlying counts and denominator. For elapsed-time measures, define the start and finish events. Without consistent definitions, a trend may reflect a changed counting method rather than a changed delivery system.

Baseline one service and follow its trend

Where possible, measure one application or service at a time, establish a baseline, and examine changes over time. DORA says its measures can apply across technology types, but warns that blending applications or teams can hide contextual differences. Deployment model, user impact, and operational risk all affect what a result means.

Use paired signals to investigate tradeoffs

  • Throughput and instability: interpret lead time and deployment frequency alongside failed deployment recovery time, change fail rate, and deployment rework rate.
  • Pre-release and post-release evidence: use test coverage as a signal of test reach, and escaped defects as evidence about problems found beyond the expected detection phase.
  • Counts and rates: raw defect or intervention counts can be misleading when deployment volume changes; rates need clear denominators and should not replace the counts that explain them.

A mismatch between measures is often more informative than a single value. For example, a team might shorten lead time while its change fail rate rises; that is a reason to inspect the delivery process and user impact, not to assume that speed itself is either good or bad.

Turn findings into improvement work

DORA recommends a practical improvement loop: establish a baseline, discuss friction, choose a significant constraint to improve, do the work, check progress, and repeat. Start with a discussion or quick check where precise data collection across multiple systems would have substantial integration costs; add instrumentation when it helps answer a consequential question.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What these metrics should not be used for

  • Do not treat them as quotas. A mandate that every application deploy multiple times daily ignores differences in service, risk, and deployment model. DORA cautions against turning metrics into targets or searching for one metric to rule a complex system.
  • Do not rank individuals or teams with raw counts. A service with more deployments has more opportunities for failures or defects to be counted. Compare a team’s measures with its own context and history, not an unqualified leaderboard.
  • Do not substitute activity for quality. The DoD guide warns that team velocity is unique to each team and should not be used to compare teams. It also warns that measuring lines of code can encourage quantity over quality.
  • Do not collapse the seven measures into a single score by default. A composite can hide tradeoffs unless the organization has a transparent, validated rationale and explains what the score means.

How AI fits into the measurement question

Google Cloud’s announcement of the 2025 DORA Report says 90% of survey respondents reported using AI at work, more than 80% believed AI increased their productivity, 30% reported little or no trust in AI-generated code, and 90% of organizations had adopted at least one platform. These are figures describing the report’s 2025 research context; they do not validate this seven-measure selection or establish that AI improved software quality. Teams considering AI-related changes still need to monitor delivery outcomes and user-facing quality in their own context.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.