What are DORA metrics?
DORA metrics are five delivery measures that show how quickly a team ships changes and how often its deployments fail or need rework. DORA stands for DevOps Research and Assessment, a research program that Google Cloud runs. Change lead time, deployment frequency, and failed deployment recovery time measure throughput. Change failure rate and deployment rework rate measure instability.
DORA's guide to the metrics calls them a way of measuring "the outcomes of the software delivery process." They are meant for one application or service at a time. Many guides still list the original four metrics. DORA's history of the metrics records that it redefined the recovery metric and later added a fifth, deployment rework rate.
The metrics follow a change through a continuous integration and delivery (CI/CD) pipeline, from its commit to its deployment in production. Change failure rate, which DORA calls change fail rate, is the metric closest to software quality, because it counts the deployments that needed immediate intervention.
How do DORA metrics work?
DORA metrics time and count what happens to changes after they are committed, over a set period. Version control records commit times, the pipeline records deployments, and incident records show each failure and its recovery. DORA's guide defines each metric this way:
- Change lead time. Change lead time, which DORA first called lead time for changes, runs from a change's commit in version control to its deployment in production. In a pull request workflow, it includes the wait for continuous integration (CI) and review.
- Deployment frequency. Deployment frequency is the number of deployments in a period, or the time between them. Under continuous delivery, a person approves each release, so the number also reflects how often the team chooses to release.
- Failed deployment recovery time. Failed deployment recovery time is how long a team takes to recover from a deployment that fails and needs immediate intervention. It replaced mean time to recovery (MTTR), which also counted failures from outside causes, e.g. a data center outage.
- Change failure rate. Change failure rate is the share of deployments that need immediate intervention, usually a rollback or a hotfix.
- Deployment rework rate. Deployment rework rate is the share of deployments that were not planned and happened because of an incident in production.
Throughput and instability are read together, because a team can raise its throughput by shipping changes that then fail. DORA's guide warns against treating any one metric as the only one that counts.
What is an example of change failure rate?
Here is an illustrative example. Acme Co. sells furniture online and deploys its web store from a CI/CD pipeline. Over four weeks, Acme's team measures its change failure rate:
- The pipeline deploys the web store to production 40 times, counting one hotfix.
- A developer asks a coding agent to "Let customers edit their delivery address during checkout." The change passes CI and deploys.
- Within 20 minutes, customers report orders with no items. The address saved, but the cart emptied.
- The team rolls back the deployment, and checkout works again 35 minutes after the deployment went out. That is the failed deployment recovery time, because the clock starts at the deployment, not at the first report.
- Later that month, a deployment breaks the order export,
acme export --format csv, and the team ships a hotfix the same day. - The team counts 2 deployments that needed immediate intervention and 1 unplanned hotfix.
The calculation divides each count by all deployments in the period:
change failure rate = failed deployments / all deployments
= 2 / 40 = 0.05 (5%)
deployment rework rate = unplanned deployments / all deployments
= 1 / 40 = 0.025 (2.5%)
The rate shows that 1 deployment in 20 failed. It does not show that the checkout tests never checked the cart after an address change. This example is simplified. A real team would also agree in advance on what counts as a failure or a deployment, e.g. whether a rollback counts as a deployment.
What changes when a coding agent writes the code?
A coding agent can write a change in minutes. Change lead time starts at the commit, so time saved before the commit does not shorten it. More agent changes mean more CI runs and reviews, and that wait can become the bottleneck for AI-written code. Throughput rises only when CI and review keep pace.
DORA's 2025 research of nearly 5,000 professionals found that AI adoption now goes with higher delivery throughput but still with lower delivery stability. DORA's explanation is that more change volume leads to instability in teams without strong controls, e.g. automated testing.
An agent's change can pass its own tests and still break working features that the task never touched. Change failure rate counts that failure only after deployment.
Change failure rate is a ratio, so it can hide a rise in failures. If Acme's next four weeks bring 80 deployments and 3 failed deployments, the rate falls, but customers get 3 broken releases instead of 2.
A team that practices agentic engineering can track the count of failed deployments next to the rate and tag deployments that carry agent-written changes. Keeping pull request size small makes each change easier to review and to roll back.
What are the limits of DORA metrics?
DORA metrics describe delivery outcomes, and they have these limits:
- They count only failures that someone detects. A deployment counts as failed only after someone finds a problem in it that needs immediate intervention, e.g. through monitoring, so the instability metrics depend on testing in production. A bug that nobody detects leaves change failure rate unchanged.
- They show that failures happened, not why. A rate cannot name the missing test or the change that caused the failure. Root cause analysis finds that cause.
- They can be gamed. DORA's guide warns that setting the metrics as goals makes teams more likely to game them, a pattern known as Goodhart's law. A team can lower its change failure rate by narrowing what it counts as a failure.
- Comparisons between teams can mislead. Teams define failures and deployments differently, and DORA's guide warns against comparing applications that differ widely. The guide sets no single target for change failure rate, so a service's rate is best judged against its own earlier results.
How are DORA metrics different from test coverage?
Test coverage is often reported as code coverage, which measures how much of the code the tests run before release. DORA metrics measure what happens to deployments after release. Coverage comes from test runs, and DORA metrics come from deployment and incident records.
The two meet at change failure rate. A suite can reach high coverage with weak checks, and the failures it misses can appear later as failed deployments. Regression testing that catches a broken feature before release keeps that failure out of the rate. Neither number shows that the software does what was asked.
FAQs
What does DORA stand for?
DORA stands for DevOps Research and Assessment. The name belongs to a research program that Google Cloud runs, and teams also use it as shorthand for the five delivery metrics that the program defines.
Do DORA metrics measure code quality?
DORA metrics do not measure code quality directly. Change failure rate counts deployments that went wrong in production, and deployment rework rate counts unplanned deployments made to fix production incidents. Both reflect quality only after release, and only for the failures that someone detects.
What replaced mean time to recovery in DORA's metrics?
Failed deployment recovery time replaced mean time to recovery in DORA's metrics. The newer metric counts only recovery from a deployment that needed immediate intervention. A failure with an outside cause, e.g. a data center outage, does not count toward it.
Can AI coding tools improve DORA metrics?
AI coding tools can improve the throughput metrics, and DORA's research found that AI adoption now goes with higher delivery throughput. The same research found that it also goes with lower delivery stability, so change failure rate can rise unless testing and other checks keep pace with the extra changes.
What is a good change failure rate?
A good change failure rate is one that improves on the same service's earlier results while the count of failed deployments does not rise. DORA's guide sets no single target, and teams count failures differently, so comparing the rate with other teams can mislead.