Quick answer: DORA metrics are the software delivery performance measures from Google's DevOps Research and Assessment programme: change lead time, deployment frequency, change failure rate and failed deployment recovery time, with deployment rework rate added in 2024. They measure throughput and stability at the team or service level. They are not a developer productivity score; pair them with the SPACE framework and developer experience surveys to see the whole picture.
Every engineering leader in Australia is being asked a version of the same question in 2026: "We have spent on AI tooling, cloud and platform teams, so what are we getting?" Boards and CFOs want a number. Engineers, rightly, do not trust numbers that reduce their work to lines of code or story points. DORA metrics are the closest thing the industry has to a shared, evidence-based answer, and the 2025 DORA report made them more relevant, not less, by showing that AI adoption raises throughput and instability at the same time.
This guide is for CTOs, heads of engineering, platform and DevOps leads and the transformation teams who report to them. It explains what the DORA metrics are, how to measure them honestly, where they break, how SPACE and DevEx complete the picture, and how to run a measurement programme that improves delivery instead of distorting it.
What are the DORA metrics?
DORA (DevOps Research and Assessment) is the research programme behind the annual State of DevOps reports and the book Accelerate. Across more than a decade of surveys it identified a small set of delivery metrics that predict organisational performance. As published on dora.dev today there are five, in two groups:
- Throughput — Metric: Change lead time · What it measures: Time from a code commit to that change running in production · How it is usually calculated: Median of (deploy timestamp minus first commit timestamp) per change
- Throughput — Metric: Deployment frequency · What it measures: How often a service deploys to production · How it is usually calculated: Count of production deployments per day/week per service
- Throughput — Metric: Failed deployment recovery time · What it measures: How long it takes to restore service after a deployment causes a failure · How it is usually calculated: Median time from a failed deployment's detection to restoration
- Instability — Metric: Change failure rate · What it measures: Share of deployments that cause a failure needing a hotfix, rollback or patch · How it is usually calculated: Failed deployments divided by total deployments
- Instability — Metric: Deployment rework rate · What it measures: Share of deployments that were unplanned and triggered by an incident in production · How it is usually calculated: Unplanned incident-driven deployments divided by total deployments
Two findings sit behind the metrics. First, speed and stability are not a trade-off: the best-performing teams are better at all five, and the worst are worse at all five. Second, the metrics describe an application or service, not an organisation and not a person. Averaging them across 400 services produces a number that nobody can act on.
DORA's performance clusters are a reasonable starting benchmark: elite teams deploy on demand (multiple times a day) with change lead times under a day, recover from failed deployments in under an hour and keep change failure rates low; low performers deploy monthly or less, with lead times measured in months. The precise cut-offs move between reports, so use them to classify your own services rather than to set targets.
How do you measure DORA metrics?
Measurement is where most programmes quietly fail, because the data lives in four different systems and each one defines "deployment" differently.
- Define "production" and "deployment" per service. A deployment is a change reaching the environment that serves real users, including dark launches behind a flag. For batch and data platforms, decide explicitly what counts.
- Instrument the pipeline, not the people. Pull deploy events from your CD tool (Argo CD, GitHub Actions, GitLab, Buildkite, Octopus), commit timestamps from version control, and incident and rollback events from your incident tooling (PagerDuty, Opsgenie, ServiceNow). Tag each incident with the deployment that caused it, otherwise change failure rate is a guess.
- Compute at service level, then roll up by team and domain. Keep the distribution, not just the mean. Lead time medians hide the 10% of changes that take six weeks because they wait on a change advisory board.
- Use tooling if you have scale. Engineering intelligence products (DX, Swarmia, Jellyfish, LinearB, Faros, the built-in DORA dashboards in GitLab and Atlassian) do the plumbing. For fewer than a dozen services, a warehouse table and a dashboard are enough.
- Publish definitions alongside the numbers. The moment a dashboard appears, someone will compare teams. Make it clear that teams own different services with different risk profiles.
In regulated Australian environments add one more step: map change failure and recovery data to the evidence your operational risk function needs. The same incident-to-deployment linkage that feeds DORA metrics is what APRA CPS 230 tolerance reporting or an agency's ICT assurance review asks for.
Why DORA metrics are not a developer productivity score
DORA metrics measure the delivery system. They say nothing about whether the right things were built, whether engineers are burning out, or whether a team's low deployment frequency is because it runs a mainframe batch that correctly deploys monthly.
Three failure modes show up repeatedly:
- Goodhart's law. Set deployment frequency as a target and teams will split changes into trivial deploys. Set change failure rate as a target and incidents stop being linked to deployments. DORA's own guidance is explicit: do not set the metrics as standalone goals.
- Individual attribution. Any attempt to compute DORA metrics per engineer destroys trust and is statistically meaningless. The unit is the service and the team.
- Context blindness. A payments core and a marketing microsite should not be on the same leaderboard.
The 2024 DORA report illustrated the point with AI: a 25% increase in AI adoption was associated with a small decrease in throughput and a larger decrease in delivery stability. Teams were producing more code faster, and the delivery system was not absorbing it. The 2025 report found throughput had turned positive while instability remained, and concluded that AI acts as an amplifier of whatever delivery capabilities an organisation already has.
The SPACE framework and developer experience
To see what DORA metrics cannot, most mature engineering organisations add two things.
SPACE (Forsgren, Storey et al., 2021) frames developer productivity across five dimensions: Satisfaction and wellbeing, Performance, Activity, Communication and collaboration, and Efficiency and flow. The rule is to pick metrics from at least three dimensions and to include at least one perceptual (survey) measure, so that system data is always balanced by what engineers report.
Developer experience (DevEx) research (Noda, Storey, Forsgren and Greiler, 2023) narrows that to three drivers: feedback loops, cognitive load and flow state. In practice that means a quarterly developer survey with a handful of stable questions (time lost to waiting, confidence in deploying, ease of finding information, satisfaction with tooling) plus a few system measures such as build time, review turnaround and time to first deploy for new starters.
A workable enterprise scorecard looks like this:
- Delivery (DORA) — System metric: Lead time, deployment frequency, change failure rate, recovery time, rework rate per service · Perceptual metric: Confidence that a change can be shipped safely today
- Flow and efficiency — System metric: Build and test time, PR review turnaround, time blocked on other teams · Perceptual metric: Time lost to friction per week (self-reported)
- Quality and reliability — System metric: SLO attainment, incident count and severity, escaped defects · Perceptual metric: Trust in the codebase and in AI-generated code
- Satisfaction — System metric: Attrition, onboarding time to first production change · Perceptual metric: Developer satisfaction, cognitive load
- Business outcomes — System metric: Features used, revenue or cost outcomes per initiative · Perceptual metric: Clarity of priorities
How do you improve DORA metrics?
The metrics are a diagnostic. The improvements come from the capabilities DORA has validated over the years, most of which are unglamorous:
- Work in small batches. Smaller changes move faster through review and testing and fail less often. This is the single highest-leverage practice and the one AI tooling most often undermines, because it makes large changes cheap to produce.
- Trunk-based development and fast, reliable test automation. Long-lived branches and flaky suites inflate lead time and hide defects.
- Continuous delivery with progressive rollout. Feature flags, canaries and automated rollback cut recovery time and make change failure rates honest.
- Loosely coupled architecture. Teams that can deploy independently deploy more often. Monolith decomposition is a DORA improvement programme, whether or not it is called one.
- An internal developer platform. Golden paths reduce the variance between teams and build the controls into the path; see our guide to platform engineering and internal developer platforms.
- Lightweight change approval. Peer review plus automated controls outperforms external change boards on both speed and stability in DORA's research.
- Generative culture. Blameless post-incident reviews and psychological safety are consistently associated with better delivery outcomes.
Are DORA metrics still relevant with AI coding assistants?
More than before. DORA's 2025 State of AI-assisted Software Development report found that around 90% of software professionals now use AI at work, that AI increases throughput, and that it also increases instability where review, testing and release practices have not caught up. The report's AI Capabilities Model names seven conditions under which AI investment improves outcomes, including working in small batches, strong version control practices, user-centred focus, healthy data and a quality internal platform.
For leaders this means two things. First, measure the delivery system before and after AI rollout, per service, with change failure rate and rework rate given equal weight to lead time. Second, treat a jump in throughput with a rise in instability as a signal to invest in review capacity, test automation and smaller batches, not as proof the tooling works. Our guide to AI coding assistants in enterprise engineering teams covers the rollout and governance side.
A 90-day plan to stand up DORA metrics
- Weeks 1-2: pick 5-10 services with different risk profiles; agree definitions of deployment, failure and production for each.
- Weeks 3-6: instrument deploy events, commits and incident linkage; stand up a dashboard with distributions, not just medians.
- Weeks 7-8: run the first DevEx survey (10 questions, anonymous, repeated quarterly).
- Weeks 9-10: review results with each team; identify the one constraint per service (CAB wait, flaky tests, manual release, review turnaround).
- Weeks 11-13: fund one improvement per team; agree that leadership reviews trends per service, never a cross-team leaderboard.
Key takeaways
- DORA metrics measure the delivery system (throughput and stability) per service; they are not a developer productivity or individual performance score.
- Instrument pipelines and incident linkage; report distributions, publish definitions and never build a cross-team leaderboard.
- Pair DORA with SPACE and a quarterly DevEx survey to capture satisfaction, flow and cognitive load.
- Improve the constraint the data reveals: small batches, test automation, progressive delivery, lightweight change approval and an internal platform.
- With AI tooling raising throughput and instability together, change failure and rework rates deserve as much attention as lead time.
Join your peers at the Clutch Engineering and DevOps Summits
Clutch Events runs free-to-attend, invite-curated, practitioner-led Engineering and DevOps summits for senior engineering, platform and DevOps leaders at Australian enterprises and government:
- Melbourne Engineering and DevOps Summit 2026 — 15 October 2026
- Sydney Engineering and DevOps Summit 2027 — 9 September 2027
- Melbourne Engineering and DevOps Summit 2027 — 7 October 2027
See all upcoming Clutch events · More guides on Clutch Events Insights