Research · Tech Industry & Practitioner

Back to sweep

Research sweep · deep · 2023 – 2026

DORA Metrics in the AI Era

How the DORA four keys (deployment frequency, lead time for changes, change failure rate, failed deployment recovery time) plus the 2024 rework-rate metric are measured and improved in practice, and how AI-assisted development has shifted them, September 2023 to September 2026: DORA State of DevOps 2023 and 2024, the 2025 State of AI-assisted Software Development and its AI Capabilities Model, the METR developer RCT, GitClear code-churn data, and the SPACE and DevEx frameworks

  • Claude Fable 5
  • tech
  • academic
  • blogs
  • financial
  • frontier
  • vc

Synthesised 2026-09-09

Narrative

DORA's own numbers on AI have moved from a startling contradiction in 2024 to a more structured "amplifier" model in 2025. The 2024 Accelerate State of DevOps report, drawing on Google Cloud survey data, found that a 25% increase in AI adoption correlated with a 7.2% reduction in delivery stability and a 1.5% reduction in delivery throughput, even as it substantially increased individual productivity, flow state and job satisfaction. DORA lead Nathen Harvey told The New Stack that engineers using more AI report better time in flow and higher individual productivity, but that this "isn't having good downstream effects when it comes to being able to deliver software fast and safe," and is "actually detrimental to their software delivery performance metrics." RedMonk's Rachel Stephens flagged the report's more consequential structural move: 2024 was the year DORA began treating recovery time as a throughput measure and introduced deployment rework rate, unplanned deployments needed to fix user-facing bugs, as a new stability signal alongside change failure rate.

The 2025 State of AI-assisted Software Development report and its companion AI Capabilities Model reframe the story around amplification rather than a simple speed-up. Google's DORA team describes AI's primary role as amplifying "the strengths of high-performing organizations and the dysfunctions of struggling ones," built from a survey of nearly 5,000 respondents, with 90% of developers now using AI daily. The Capabilities Model itself was built through 78 in-depth interviews followed by validated survey questions tested against the same near-5,000-person sample, yielding seven capabilities: a clear and communicated AI stance, healthy data ecosystems, AI-accessible internal data, strong version control practices, working in small batches, user-centric focus, and quality internal platforms. Several of these (small batches, version control discipline, platform quality) map directly onto pre-existing DORA capabilities that predict the four keys, reinforcing that AI's benefit is conditional on fundamentals already being in place rather than a substitute for them.

Independent evidence complicates the productivity story further. METR's July 2025 randomised controlled trial, run with 16 experienced open-source developers completing 246 real tasks in mature repositories they had worked on for an average of five years, found that allowing AI tools (mainly Cursor Pro with Claude 3.5/3.7 Sonnet) made developers 19% slower, even though those same developers had forecast a 24% speedup beforehand and estimated after the fact that AI had made them roughly 20% faster. This is one controlled study in one setting (February to June 2025 frontier tools) and MEtr itself now labels the finding as historical, but the divergence between measured time and self-reported speed is the sharpest empirical break from vendor productivity claims found in this sweep. GitClear's commit-level telemetry, spanning over 211 million and, in its 2026 update, 600 million analysed lines of code, corroborates a related but distinct concern: 2024 was the first year on record where copy-pasted code exceeded "moved" (refactored) code within commits, duplicated code blocks rose roughly eightfold to tenfold over two years, and code churn (lines revised within two weeks of authoring) climbed from roughly 3-5% in 2020-2021 to 5.7-7.9% by 2024.

Measurement practice around the four/five keys remains uneven outside DORA's own survey instrument. DORA's canonical definitions (deployment frequency, change lead time from commit to production, change failure rate as percentage of deployments needing hotfix or rollback, and failed deployment recovery time) come from an annual self-report survey reaching over 36,000 to 39,000 professionals across the programme's history, not from a single standardised telemetry pipeline, which is why practitioner tools such as Swarmia, the CD Foundation and IBM's engineering guides describe teams instrumenting the same metrics through CI/CD and VCS data with materially different operational choices (for example whether rollbacks count as new deployments). The "elite performer" tier itself is a moving target: Google's own 2020-era Four Keys posts note that the 2022 report's cluster analysis stopped detecting an elite cluster at all, before elite reappeared in the 2023 and 2024 waves, meaning any citation of "elite" thresholds needs to be tied to a specific report year rather than treated as a stable benchmark. On frameworks, DORA sits alongside SPACE (2021, five dimensions of developer productivity) and the DevEx framework published in ACM Queue in April 2023 by Abi Noda, Margaret-Anne Storey, Nicole Forsgren and Michaela Greiler, which distills productivity into feedback loops, cognitive load and flow state; DX Core 4 (2024) is a further consolidation by overlapping authors, and Thoughtworks' Technology Radar entry on DORA metrics argues that in the AI era, "measuring productivity by lines of code generated by AI is misleading" and that real improvement must show up in delivery flow and stability, not raw output.

Thoughtworks' Technology Radar tracks the qualitative texture behind DORA's numbers. Volume 32 (April 2025) and Volume 33 (November 2025) both flag AI-related antipatterns, including complacency with AI-generated code and AI-accelerated shadow IT, alongside genuine gains such as AI helping teams understand legacy codebases; CTO Rachel Laycock noted that "vibe coding" as an approach had "practically disappeared" by late 2025 in favour of more disciplined context engineering. This trajectory, disciplined small-batch practice paired with AI rather than AI as a wholesale replacement for engineering judgement, is consistent with what both DORA's Capabilities Model and GitClear's churn data independently point toward: AI adoption without attention to batch size, version control discipline, and platform quality tends to show up as instability and rework rather than net throughput gains.


Sources

ID Title Outlet Date Significance
p1 DORA: AI boosting productivity, hindering delivery Substack (practitioner) 2024 Highlights the counterintuitive 2024 finding that AI improved code quality/review speed indicators while lowering lead time, throughput and change failure rate.
p2 Introducing the DORA AI Capabilities Model: 7 keys to succeeding in AI-assisted software development Google Research 2025-09 Google Research write-up describing the three-phase methodology (78 interviews, validated survey, ~5,000-respondent analysis) behind the capabilities model.
p3 From adoption to impact: Putting the DORA AI Capabilities Model to work Google Cloud Blog 2025-12 DORA lead Nathen Harvey's own framing of AI as an amplifier, citing 90% daily AI adoption and cluster analysis showing outcome variation across teams.
p4 The 2024 DevOps Performance Clusters Octopus Deploy blog 2024-11 Explains how DORA's elite/high/medium/low clusters are derived statistically each year and cautions against treating any single year's thresholds as fixed benchmarks.
p5 The DORA 4 key metrics become 5 Continuous Delivery Foundation 2025-10 CD Foundation explainer detailing exactly how rework rate is calculated and why recovery time moved from stability to throughput.
p6 DORA | A history of DORA's software delivery metrics DORA / Google Cloud 2026-01 DORA's own chronology of metric changes since 2014, including the 2018 addition of availability and the 2024 addition of deployment rework rate.
p7 DORA metrics | Technology Radar Thoughtworks Technology Radar 2026-04 Thoughtworks' practitioner take arguing that in the AI era, lines-of-code output is a misleading productivity signal versus delivery flow and stability.
p8 Thoughtworks Technology Radar Highlights The Rapid Evolution of AI Assistance in 2025 Thoughtworks 2025-11 CTO Rachel Laycock's account of AI antipatterns (shadow IT, complacency with AI code) and the shift from vibe coding to context engineering across 2025.
p9 From vibe coding to context engineering: 2025 in software development Thoughtworks 2026-06 Thoughtworks practitioner reflection connecting context engineering practice to reduced rework and better AI code reliability.
p10 How to Measure Developer Productivity: DORA vs SPACE vs GSM ShiftMag 2026-03 Situates DORA against the SPACE (2021) and DevEx (April 2023, ACM Queue) frameworks by original named authors including Nicole Forsgren and Margaret-Anne Storey.
p11 A new way to measure developer productivity – from the creators of DORA and SPACE The Pragmatic Engineer (Gergely Orosz) 2024-05 Interview with DevEx framework authors explaining the three-dimension (feedback loops, cognitive load, flow state) model and its lineage from DORA and SPACE.
p12 DORA 2025 AI Capabilities: How to Boost AI Impact kt.team December 9, 2025 Retrieved by this lane's web search.
p13 2025 AI Capabilities Model meriroos.ee Retrieved by this lane's web search.
p14 Quantifying the Expectation-Realisation Gap for Agentic AI Systems arxiv.org Retrieved by this lane's web search.
p15 DORA Metrics Explained (2026): Four Keys & Benchmarks taskade.com June 9, 2026 Retrieved by this lane's web search.
p16 Use Four Keys metrics like change failure rate to measure your DevOps performance | Google Cloud Blog cloud.google.com September 22, 2020 Retrieved by this lane's web search.
p17 DORA Metrics: Measure Open DevOps Success with 4 Key Indicators waydev.co December 16, 2025 Retrieved by this lane's web search.
p18 What Elite Performing Teams Measure | An Introduction to DORA Metrics | Lean TECHniques leantechniques.com March 2, 2026 Retrieved by this lane's web search.
p19 DORA metrics: the complete guide to measuring DevOps performance in the AI era getdx.com July 13, 2026 Retrieved by this lane's web search.
p20 DORA Metrics Benchmarks 2026: Where Does Your Team Stand? | Gitrecap Blog gitrecap.com April 9, 2026 Retrieved by this lane's web search.
p21 Jump to Content cloud.google.com Retrieved by this lane's web search.
p22 SPACE Metrics Framework for Developers Explained (2025 Edition) | LinearB Blog linearb.io April 18, 2025 Retrieved by this lane's web search.
p23 DevOps Metrics: DORA Metrics, SPACE Framework, DevEx - Travis CI travis-ci.com May 30, 2025 Retrieved by this lane's web search.
p24 Unlocking Engineering Excellence: DORA, SPACE, DevEX, and Beyond | by Soulaiman Ghanem | Code Factory Berlin | Medium medium.com February 1, 2025 Retrieved by this lane's web search.
p25 Harnessing the Potential of Gen-AI Coding Assistants in Public Sector Software Development arxiv.org Retrieved by this lane's web search.

We use analytics cookies to understand site usage and improve the service. We do not use marketing cookies.