Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

DORA metrics show whether a software service is being delivered faster without becoming less stable. The current DORA model has five measures: change lead time, deployment frequency, failed deployment recovery time, change fail rate, and deployment rework rate. Track them for one service at a time, using consistent event definitions, and read delivery speed alongside production disruption—not as a score for individual developers.

What are DORA metrics?

DORA metrics are measures of software delivery performance. They describe outcomes for a team’s application or service, rather than how hard any individual developer worked. DORA describes them as leading indicators for organizational performance and employee well-being, and lagging indicators for software development and delivery practices (DORA metrics guide).

The current model has five metrics. Three describe throughput—the pace of delivery and recovery—while two describe instability associated with production changes. They are most useful together: a faster release cadence is not an improvement if failures or unplanned repair work rise sharply.

Metric What it measures Practical calculation
Change lead time Elapsed time from a code commit until that change is successfully deployed to production. For each production deployment, subtract the relevant commit timestamp from its successful production deployment timestamp. Report a defined distribution or summary over a consistent period.
Deployment frequency How often the service is deployed to production. Count qualifying production deployments in the period, or report the interval between deployments.
Failed deployment recovery time Time needed to recover from a failed deployment that requires immediate intervention. For each qualifying failure, subtract the failure or intervention start time from the time the service is recovered. Define the event boundaries before measuring.
Change fail rate The proportion of production deployments that require immediate intervention, such as a rollback or hotfix. Number of deployments requiring immediate intervention divided by all qualifying production deployments in the same period, multiplied by 100.
Deployment rework rate The proportion of deployments that are unplanned work caused by a production incident. Number of unplanned, incident-driven deployments divided by all qualifying deployments in the same period, multiplied by 100.

For each measure, document which events qualify and how you aggregate the observations. Averages alone can conceal long delays or a small number of severe failures, so select and retain a reporting method that fits the service and use it consistently.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How the current five metrics differ from the original Four Keys

The original Four Keys model used deployment frequency, lead time for changes, change failure rate, and time to restore service. Google Cloud’s historical Four Keys article says DORA added reliability in 2021; the current DORA guide instead presents five metrics, including failed deployment recovery time and deployment rework rate (Google Cloud’s Four Keys article; current DORA metrics guide).

Terminology matters when reading older dashboards or reports. “Time to restore service” and “failed deployment recovery time” are related recovery concepts, but do not assume every organization defines their start and end events identically. Likewise, reliability is an operational concern—whether a service meets or exceeds its reliability targets—not simply another name for deployment speed. The 2021 State of DevOps report discussed reliability alongside the four delivery measures (Google Cloud’s 2021 report summary).

How to calculate and implement DORA metrics

The calculations depend on trustworthy event data. Before comparing results, define the production boundary, the events counted as failures, and how changes are linked to deployments. Then preserve the underlying event history so that a changed definition does not silently rewrite past measurements.

  1. Define the service and its production boundary. Decide what counts as a production deployment for that application, including how you treat staged rollouts, regional releases, and automated promotions. Write down what constitutes a failed deployment and an intervention.
  2. Collect timestamped events. Capture commits from source control; deployment successes and failures from the CI/CD system; rollback and hotfix events; and incident and recovery timestamps from the incident system. Record a service identifier and deployment identifier wherever possible.
  3. Join related events. Associate commits with the deployment that delivered them and incidents or corrective deployments with the affected service and deployment. Keep raw events as well as derived metrics, enabling audits when definitions change.
  4. Calculate the five measures over a fixed window. Use the same period boundaries and rules for the numerator and denominator of each rate. Keep the aggregation method stable when reviewing a trend.
  5. Review throughput and instability together. Look for changes in lead time and deployment cadence alongside failures, recovery duration, and incident-driven rework. Investigate the process or system conditions behind a change, then test an improvement and monitor its effect.

Google’s Four Keys reference implementation illustrates one way to assemble the data: a generalized ETL pipeline receives webhook events, parses changes, deployments, and incidents, and loads them into BigQuery for dashboards. Google says tools capable of emitting an HTTP request can be integrated (Four Keys implementation overview). The underlying principle is tool-neutral: source control, CI/CD, and incident systems must emit events that can be linked reliably.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to interpret results without creating a misleading score

Measure one application or service at a time, and compare its trend in context. DORA cautions that the metrics are best suited to an individual application or service; combining unlike teams, architectures, or services can obscure meaningful differences (DORA metrics guide; DORA capabilities guidance).

When comparing services or teams, check whether they share the same production boundary, deployment and incident definitions, reporting window, and aggregation method. Also account for architecture and reliability targets. A service with a different release model or operational role may not be fairly compared with another through a single raw number.

  • Use trends for improvement. A service’s movement over time is generally more informative than ranking teams against one another.
  • Keep definitions stable within a reporting period. If instrumentation or event rules change, record when and why; otherwise an apparent improvement may reflect a measurement change.
  • Pair speed with stability. Raising deployment frequency alone can reward a behavior that increases failures, recovery work, or incident-driven releases.
  • Do not use the figures as individual quotas. DORA metrics measure delivery outcomes at the service or team level, not developer productivity or effort.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What do historical DORA benchmarks say?

Published performance bands can provide orientation, but they are snapshots of survey cohorts, not universal thresholds for every service. The figures below come from Google Cloud/DORA’s 2021 report, which surveyed more than 32,000 professionals (2021 report summary).

Measure Elite performers in the 2021 report Low performers in the 2021 report
Deployments About 1,460 per year About 1.5 per year
Change lead time Under one hour More than six months
Time to restore service Under one hour More than six months
Change failure rate 0%–15% band; 7.5% midpoint estimate 16%–30% band; 23% midpoint estimate

In that 2021 comparison, elite performers deployed approximately 973 times as frequently as low performers, based on the reported annual deployment figures. These numbers are tied to the report’s categories and terminology; they are not targets to impose on a service with different constraints. The 2023 State of DevOps summary reports a survey cohort of more than 36,000 professionals and says year-over-year comparison is more meaningful than comparison with other companies (2023 report summary).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.