Research · DORA Metrics in the AI Era

Back to research

Research sweep · deep · 2023 – 2026

DORA Metrics in the AI Era

How the DORA four keys (deployment frequency, lead time for changes, change failure rate, failed deployment recovery time) plus the 2024 rework-rate metric are measured and improved in practice, and how AI-assisted development has shifted them, September 2023 to September 2026: DORA State of DevOps 2023 and 2024, the 2025 State of AI-assisted Software Development and its AI Capabilities Model, the METR developer RCT, GitClear code-churn data, and the SPACE and DevEx frameworks

Explore the research lanes ↗

Synthesised 2026-09-09

DORA in the AI era: the metrics held, the practices got louder

Overview

The four keys survived AI. Everything around them shifted. Between September 2023 and September 2026 DORA renamed one metric (time to restore became failed deployment recovery time, redefined in 2023 to isolate failures caused by a production change), reclassified another (recovery time moved from stability to throughput in 2024), added a fifth (rework rate, unplanned deployments fixing user-facing bugs), and in 2025 rebranded the whole programme as the State of AI-assisted Software Development, retiring low-to-elite performance tiers in favour of seven team archetypes. Sources: DORA / Google Cloud (2026) (); Continuous Delivery Foundation (2025) (); redmonk.com (2024) (); RedMonk (2025) ()

The defining shift of the past eighteen months is the collapse of the simple speed-up narrative. DORA 2024 measured a 25% increase in AI adoption as correlating with a 7.2% reduction in delivery stability and a 1.5% reduction in throughput, even as individual productivity and flow rose. METR's randomised trial then found experienced developers 19% slower with AI while believing themselves 20% faster. By 2025 DORA had replaced trade-off language with an amplifier thesis: AI magnifies the strengths of well-run organisations and the dysfunctions of struggling ones. Sources: dora.dev (n.d.) (); metr.org (2025) (); Google Cloud Blog (2025) ()

For a leader deciding what to change on Monday, the practical message is conservative. The capabilities that predicted the four keys before AI (small batches, version control discipline, continuous testing, platform quality) are the same ones DORA's 2025 AI Capabilities Model found determine whether AI helps or harms delivery. AI raises the price of skipping them, because it makes it trivially easy to enlarge batch size, duplicate code, and outrun review capacity. Sources: Google Research (2025) (); Substack (practitioner) (2024) ()

Timeline

Key milestones, H2 2023 to H1 2026
H2 2023
  • DORA redefines recovery metric as failed deployment recovery time, isolating change-caused failures
H1 2024
  • GitHub and Accenture publish enterprise Copilot ROI figures, setting the vendor benchmark for gains
H2 2024
  • DORA 2024 measures AI adoption reducing throughput and stability
  • Rework rate added as fifth metric, recovery time reclassified as throughput
H1 2025
  • GitClear reports copy-pasted code exceeding refactored code for the first time
  • Cursor maker Anysphere valued at 9.9 billion dollars, capital fully committed to the acceleration thesis
H2 2025
  • METR RCT finds 19% slowdown against a forecast 24% speedup
  • DORA rebrands as State of AI-assisted Software Development, ships the seven-capability AI model and seven team archetypes
  • Gartner publishes first Magic Quadrant for AI code assistants
H1 2026
  • METR abandons its follow-up study after developers refuse the no-AI arm
  • GitClear extends telemetry, block duplication up 81% on the 2023 baseline
  • Vibe coding gives way to disciplined context engineering per Thoughtworks

Key Findings

Throughput and stability are now framed as companions, not a trade-off, and the metric groupings changed to enforce that. DORA 2024 formally paired deployment frequency and lead time as throughput and change failure rate and recovery time as stability, then moved recovery time into throughput and treated change failure rate as a proxy for rework, adding deployment rework rate to measure it directly. RedMonk's Rachel Stephens flagged this as the report's most consequential structural move. The decade-long finding behind it is that high performers achieve both at once, so any local improvement that trades one for the other is a mis-optimisation. Sources: redmonk.com (2024) (); Continuous Delivery Foundation (2025) (); DORA / Google Cloud (2026) ()

The strongest levers per metric are boring and repeated across report years. Lead time responds to batch size and trunk-based version control practice: DORA's explanation for the 2024 AI penalty was precisely that AI enlarges changesets, and "DORA's data has consistently shown that larger changesets introduce risk". Change failure rate and rework respond to continuous testing and review discipline. Recovery time responds to version control hygiene and rollback capability, which is why "strong version control practices" appears verbatim in the 2025 AI Capabilities Model. The Monday action is mechanical: cap PR size, keep AI-generated work on short-lived branches, and refuse to let scaffolding volume masquerade as throughput. Sources: thenewstack.io (2024) (); Google Research (2025) (); aviator.co (2026) ()

DORA 2024's headline was a measured contradiction, and it replicated only partially. With roughly 39,000 respondents and about 3,000 answering AI questions, the 2024 Bayesian model tied a 25% AI adoption increase to a 7.2% stability drop, a 1.5% throughput drop, and a 2.1% individual productivity gain. The 2025 wave (near 5,000 respondents, over 100 hours of interviews) flipped the throughput sign positive while instability persisted; DORA's own language was that AI "not only fails to fix instability, it is currently associated with increasing instability". The stability finding has now held across two survey waves. The throughput finding has not, so treat it as unsettled. Sources: dora.dev (n.d.) (); RedMonk (2025) (); Research-Driven Engineering Leadership (Substack) (2025) ()

METR's RCT is the field's only controlled evidence, and it broke the self-report link. Sixteen experienced open-source developers, 246 real tasks in repositories they had maintained for an average of five years, mostly Cursor Pro with Claude 3.5/3.7 Sonnet: measured completion time rose 19% (95% CI roughly -40% to -2%) against a forecast 24% speedup and a post-hoc estimate of 20% faster. Independent readings from LessWrong and others argue the result may not generalise beyond experts in familiar code, and METR itself now labels it historical, stating in February 2026 that it believes developers are more sped up by early-2026 AI than the study measured. The durable lesson is narrower and more useful: developer self-report is an unreliable instrument for AI productivity, so measure delivery outcomes instead. Sources: arXiv / METR (2025) (); LessWrong (2025) (); METR (Substack) (2026) (); Reuters (via Investing.com) (2025) ()

GitClear's telemetry corroborates the instability side, not the productivity side. Across 211 million changed lines (2020 to 2024), copy-pasted code exceeded moved (refactored) code for the first time in 2024, duplicated blocks rose roughly eightfold, and two-week churn climbed from about 3.1% in 2020 to 5.7% in 2024, with moved code falling from about 25% of changed lines in 2021 to under 10%. The 2026 follow-up found block duplication up 81% on the 2023 baseline, and, tellingly, that heavy AI users' 4 to 10x output advantage over non-users mostly predates their AI adoption, with a more modest 25% velocity gain against their own past output. This is vendor telemetry from one company's customer base, but it is the only commit-level dataset at this scale and it points the same direction as DORA's stability finding. Sources: DevClass (2025) (); particula.tech (2026) (); DevOps.com (citing Wall Street Journal) (2025) ()

The 2025 AI Capabilities Model is mostly the old capability list with an AI prefix, and that is its point. Built from 78 interviews and validated against the near-5,000 sample, its seven capabilities are a clear and communicated AI stance, healthy data ecosystems, AI-accessible internal data, strong version control practices, working in small batches, user-centric focus, and quality internal platforms, selected from fifteen candidates for showing a statistically significant interaction with outcomes. Three of the seven (small batches, version control, platforms) are pre-existing DORA capabilities. The genuinely new ones concern making internal context available to AI tools and telling people what they are allowed to do with them, both cheap to implement and, per DORA, both currently rare. Sources: Google Research (2025) (); services.google.com (n.d.) (); Google Cloud Blog (2025) ()

Platform engineering shows the same bifurcation as AI itself. DORA 2024 found internal platforms lifted individual productivity by roughly 8% and organisational performance by roughly 6% while reducing throughput by about 8% and stability by about 14%. The pattern (individual gains, delivery losses) mirrors the AI finding, suggesting a general failure mode: tools that make individuals feel faster while delivery discipline erodes. A platform is a lever only when paired with the batch-size and testing practices it is supposed to encode. Sources: thenewstack.io (2024) (); redmonk.com (2024) ()

SPACE, DevEx and DX Core 4 now bracket DORA rather than compete with it. The DevEx framework (Noda, Storey, Forsgren, Greiler, ACM Queue 2023) explicitly measures the experience "before and after" the commit-to-production window DORA covers, via feedback loops, cognitive load and flow. DX Core 4 (2024) consolidates the lot. GetDX's practitioner argument is that the four keys under-measure AI-assisted teams because deployment frequency inflates with AI scaffolding while review overhead for AI code often exceeds that for human code, so review load and cognitive load belong on the dashboard next to the five keys. Sources: The Pragmatic Engineer (Gergely Orosz) (2024) (); GetDX Newsletter (Substack) (2024) (); Unblocked (company blog) (2026) (); Jens Oliver Meiert (personal blog) (2025) ()

A good team in 2026 looks like DORA's "harmonious high-achievers" archetype, and its numbers are behavioural, not threshold-based. Ninety per cent of respondents use AI daily at a median two hours, yet 30% report little or no trust in AI output, and the good teams treat that distrust as process: small batches enforced on AI-generated changes, rework rate tracked as a first-class stability signal, internal data wired into tools, and an explicit organisational AI stance. Thoughtworks' Rachel Laycock reports vibe coding "practically disappeared" by late 2025 in favour of context engineering, which is the same conclusion Simon Willison reached from practice: AI multiplies existing discipline in testing, planning, review and documentation. Elite multipliers such as 127x lead time still circulate from the 2024 report, but the tier structure behind them no longer exists, so quote them with a year attached or not at all. Sources: RedMonk (2025) (); cloud.google.com (n.d.) (); Thoughtworks (2026) (); Simon Willison's Weblog (2025) ()

Evidence & Data

The load-bearing numbers, by source type. DORA survey (organisational, correlational): 2024, 25% more AI adoption maps to -7.2% stability, -1.5% throughput, +2.1% individual productivity, on ~39,000 respondents; 2025, throughput turns positive, instability persists, 90% daily AI use, 30% low trust, on ~5,000 respondents. Controlled trial (individual, causal): METR, 16 developers, 246 tasks, +19% completion time against forecast -24%. Commit telemetry (codebase, observational): GitClear, churn 3.1% to 5.7% (2020 to 2024), duplication up roughly eightfold in 2024, up 81% on 2023 by 2026, cloned code linked to 15 to 50% more defects. Sources: dora.dev (n.d.) (); Google Blog (2025) (); metr.org (2025) (); particula.tech (2026) ()

These are different units of analysis and should never be merged. Cui and colleagues' pooled field experiments across Microsoft, Accenture and a Fortune 100 firm (4,867 developers) found a 26.08% increase in completed tasks, with larger gains for juniors, while Peng's original Copilot RCT found 55.8% on a single scripted task. Both measure task output, not delivery. Bain puts realised team-level gains at 10 to 15%, noting coding is only 25 to 35% of the lifecycle; Gartner forecasts 30% "through 2028" as projection, not measurement; McKinsey's Jellyfish-partnered claim of over 110% gains at 80 to 100% adoption is vendor-sourced and unreplicated. Sources: cmustatistics.github.io (2025) (); Bain & Company (2025) (); Gartner (2025) (); McKinsey & Company (2025) ()

The market priced the bull case regardless: Anysphere at 9.9 billion dollars in June 2025 past 500 million ARR, past 2 billion ARR by early 2026, with talks at a 50 billion valuation. Sources: TechCrunch (citing Bloomberg) (2025) (); TechCrunch (citing Bloomberg) (2026) ()

Signals & Tensions

Measured versus asserted. DORA's Bayesian survey model, METR's RCT and GitClear's telemetry form the measured tier. Vendor telemetry (Faros, Jellyfish), Copilot ROI studies and Magic Quadrant placements form the asserted tier, and consistently report larger, more favourable numbers. The gap between the two tiers runs 30 to 50 percentage points. Sources: GitHub Blog (2024) (); faros.ai (2026) (); The Register (2025) ()

Instrumentation drift. DORA's canonical definitions come from self-report, while tool vendors instrument the same names through CI/CD data with materially different choices, such as whether rollbacks count as deployments or lead time starts at ticket rather than commit. Both errors flatter the numbers, and the survey-versus-telemetry gap is largely unquantified. Sources: getdx.com (2026) (); dynatrace.com (2025) ()

The METR refusal problem is itself a finding. The follow-up died because developers would not accept randomisation into a no-AI arm, which biases future estimates downward and means the cleanest study design in the field may be unrepeatable. Sources: METR (Substack) (2026) (); METR (2026) ()

Capability trend versus delivery trend. METR's time-horizon evaluations show autonomous coding capability climbing steadily (GPT-5 at roughly 2h15m at 50% reliability against o3's 1h30m) while team-level stability telemetry worsens with adoption. Both can be true; the bottleneck has moved to integration, review and rework. Sources: METR (2025) (); metr.org (2025) ()

Underreported: skill formation. A 2026 randomised study found AI use impaired conceptual understanding and debugging without average efficiency gains, a cost no delivery metric currently captures. Sources: ainewstoday.substack.com (n.d.) ()

Open Questions

Whether the 2025 throughput reversal is a real learning effect or an artefact of a changed sample and instrument, since DORA's waves are not a panel and question wording shifted with the rebrand. Sources: DORA (2025) ()

Whether rework rate and the AI-era proposals (AI-attributable rework, acceptance rate, trust, review load) survive a second survey wave, or join the churn of retired constructs like the elite tier. Sources: RedMonk (2025) (); GetDX Newsletter (Substack) (2025) ()

Whether GitClear's duplication and churn trends replicate outside its own customer base; the first independent RCTs on AI-assisted maintainability are only now appearing on arXiv. Sources: DevClass (2025) ()

How much of the individual-versus-organisational gap is review bottleneck, measurable as review load, versus genuine quality decay, measurable as defects, since the interventions differ. Sources: Unblocked (company blog) (2026) ()

Whether a stable elite profile for AI-assisted teams exists at all, given the archetypes are one wave old and the cluster analysis that once defined elite has already failed once, in 2022. Sources: cloud.google.com (2020) (); Octopus Deploy blog (2024) ()

Whether any credible RCT of AI-era delivery is still possible now that developers refuse control conditions, or whether the field is permanently limited to observational telemetry and survey correlation. Sources: METR (Substack) (2026) ()

The Monday takeaway holds either way: the five keys still measure the right things, AI has not changed which practices move them, and the teams winning in 2026 are the ones that got stricter about small batches and review exactly when their tools made it easiest to abandon both.


Sources

Summary: ↑ Back to summary


Tech Industry & Practitioner

ID Title Outlet Date Significance
p1 DORA: AI boosting productivity, hindering delivery Substack (practitioner) 2024 Highlights the counterintuitive 2024 finding that AI improved code quality/review speed indicators while lowering lead time, throughput and change failure rate.
p2 Introducing the DORA AI Capabilities Model: 7 keys to succeeding in AI-assisted software development Google Research 2025-09 Google Research write-up describing the three-phase methodology (78 interviews, validated survey, ~5,000-respondent analysis) behind the capabilities model.
p3 From adoption to impact: Putting the DORA AI Capabilities Model to work Google Cloud Blog 2025-12 DORA lead Nathen Harvey's own framing of AI as an amplifier, citing 90% daily AI adoption and cluster analysis showing outcome variation across teams.
p4 The 2024 DevOps Performance Clusters Octopus Deploy blog 2024-11 Explains how DORA's elite/high/medium/low clusters are derived statistically each year and cautions against treating any single year's thresholds as fixed benchmarks.
p5 The DORA 4 key metrics become 5 Continuous Delivery Foundation 2025-10 CD Foundation explainer detailing exactly how rework rate is calculated and why recovery time moved from stability to throughput.
p6 DORA | A history of DORA's software delivery metrics DORA / Google Cloud 2026-01 DORA's own chronology of metric changes since 2014, including the 2018 addition of availability and the 2024 addition of deployment rework rate.
p7 DORA metrics | Technology Radar Thoughtworks Technology Radar 2026-04 Thoughtworks' practitioner take arguing that in the AI era, lines-of-code output is a misleading productivity signal versus delivery flow and stability.
p8 Thoughtworks Technology Radar Highlights The Rapid Evolution of AI Assistance in 2025 Thoughtworks 2025-11 CTO Rachel Laycock's account of AI antipatterns (shadow IT, complacency with AI code) and the shift from vibe coding to context engineering across 2025.
p9 From vibe coding to context engineering: 2025 in software development Thoughtworks 2026-06 Thoughtworks practitioner reflection connecting context engineering practice to reduced rework and better AI code reliability.
p10 How to Measure Developer Productivity: DORA vs SPACE vs GSM ShiftMag 2026-03 Situates DORA against the SPACE (2021) and DevEx (April 2023, ACM Queue) frameworks by original named authors including Nicole Forsgren and Margaret-Anne Storey.
p11 A new way to measure developer productivity – from the creators of DORA and SPACE The Pragmatic Engineer (Gergely Orosz) 2024-05 Interview with DevEx framework authors explaining the three-dimension (feedback loops, cognitive load, flow state) model and its lineage from DORA and SPACE.
p12 DORA 2025 AI Capabilities: How to Boost AI Impact kt.team December 9, 2025 Retrieved by this lane's web search.
p13 2025 AI Capabilities Model meriroos.ee Retrieved by this lane's web search.
p14 Quantifying the Expectation-Realisation Gap for Agentic AI Systems arxiv.org Retrieved by this lane's web search.
p15 DORA Metrics Explained (2026): Four Keys & Benchmarks taskade.com June 9, 2026 Retrieved by this lane's web search.
p16 Use Four Keys metrics like change failure rate to measure your DevOps performance | Google Cloud Blog cloud.google.com September 22, 2020 Retrieved by this lane's web search.
p17 DORA Metrics: Measure Open DevOps Success with 4 Key Indicators waydev.co December 16, 2025 Retrieved by this lane's web search.
p18 What Elite Performing Teams Measure | An Introduction to DORA Metrics | Lean TECHniques leantechniques.com March 2, 2026 Retrieved by this lane's web search.
p19 DORA metrics: the complete guide to measuring DevOps performance in the AI era getdx.com July 13, 2026 Retrieved by this lane's web search.
p20 DORA Metrics Benchmarks 2026: Where Does Your Team Stand? | Gitrecap Blog gitrecap.com April 9, 2026 Retrieved by this lane's web search.
p21 Jump to Content cloud.google.com Retrieved by this lane's web search.
p22 SPACE Metrics Framework for Developers Explained (2025 Edition) | LinearB Blog linearb.io April 18, 2025 Retrieved by this lane's web search.
p23 DevOps Metrics: DORA Metrics, SPACE Framework, DevEx - Travis CI travis-ci.com May 30, 2025 Retrieved by this lane's web search.
p24 Unlocking Engineering Excellence: DORA, SPACE, DevEX, and Beyond | by Soulaiman Ghanem | Code Factory Berlin | Medium medium.com February 1, 2025 Retrieved by this lane's web search.
p25 Harnessing the Potential of Gen-AI Coding Assistants in Public Sector Software Development arxiv.org Retrieved by this lane's web search.

Academic & arXiv

ID Title Outlet Date Significance
a1 Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity arXiv / METR 2025-07 placeholder
a2 Measuring the Impact of Early-2025 AI on Experienced Open ... metr.org Retrieved by this lane's web search.
a3 AI Coding Tools Made Developers 19% Slower: METR Study | Let's Data Science letsdatascience.com February 20, 2026 Retrieved by this lane's web search.
a4 Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity - METR metr.org July 10, 2025 Retrieved by this lane's web search.
a5 Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity simonwillison.net July 12, 2025 Retrieved by this lane's web search.
a6 Research - METR metr.org Retrieved by this lane's web search.
a7 Randomized Controlled Trials for Conditional Access Optimization Agent arxiv.org Retrieved by this lane's web search.
a8 Randomized Controlled Trials for Phishing Triage Agent arxiv.org Retrieved by this lane's web search.
a9 Effect of AI on programming productivity – CMU S&DS Data Repository cmustatistics.github.io November 5, 2025 Retrieved by this lane's web search.
a10 A randomized trial by METR found that experienced developers completed real coding tasks 19% slower when allowed to use AI tools - yet afterwards, they estimated on average that AI had made them 20% faster. - ScienceBlog.com scienceblog.com July 20, 2026 Retrieved by this lane's web search.
a11 DORA 2024: AI and Platform Engineering Fall Short - The New Stack thenewstack.io November 11, 2024 Retrieved by this lane's web search.
a12 DORA | Accelerate State of DevOps Report 2024 dora.dev Retrieved by this lane's web search.
a13 DORA Report 2024 – A Look at Throughput and Stability redmonk.com December 4, 2024 Retrieved by this lane's web search.
a14 2024 DORA Report - by Abi Noda and Laura Tacho newsletter.getdx.com October 30, 2024 Retrieved by this lane's web search.
a15 Highlights from the 2024 DORA State of DevOps Report getdx.com Retrieved by this lane's web search.
a16 What the 2024 DORA Report Reveals About AI and DevOps | Mezmo mezmo.com Retrieved by this lane's web search.
a17 State of DevOps 2024: DORA Report Findings - Interactive Report libertify.com April 20, 2026 Retrieved by this lane's web search.
a18 Unveiling the 2024 DORA Accelerate State of DevOps Report: Key Insights for Your Team - Nimble Evolution us.nimbleevolution.com November 20, 2024 Retrieved by this lane's web search.
a19 DORA's State of DevOps in 2024 dlsthoughts.substack.com Retrieved by this lane's web search.
a20 2024 dora report infoq.com Retrieved by this lane's web search.
a21 2025 Dora Ai Capabilities Model | PDF | Artificial Intelligence | Intelligence (AI) & Semantics scribd.com Retrieved by this lane's web search.
a22 DORA | DORA Publications dora.dev Retrieved by this lane's web search.
a23 DORA AI Capabilities Model: AI makes good teams better blog.doubleslash.de March 25, 2026 Retrieved by this lane's web search.
a24 DORA | Download the DORA AI Capabilities Model report dora.dev November 25, 2025 Retrieved by this lane's web search.
a25 2025 DORA AI Capabilities Model services.google.com Retrieved by this lane's web search.

Blogs & Independent Thinkers

ID Title Outlet Date Significance
b1 DORA 2025: Measuring Software Delivery After AI RedMonk 2025-12 Follow-up independent analysis tracking the 2025 reversal in AI-throughput correlation and continued instability correlation, with direct quotes from the DORA report
b2 On METR's AI Coding RCT LessWrong 2025-07 LessWrong community analysis testing the METR RCT's generalisability, weighing the experience-with-tool hypothesis against the raw slowdown finding
b3 We are Changing our Developer Productivity Experiment Design METR (Substack) 2026-02 METR's own Substack admission that a 2026 follow-up study produced an unreliable signal due to participation bias, an important caveat on replication
b4 AI coders think they're 20% faster - but they're actually 19% slower Pivot to AI 2025-07 Skeptical independent blog treatment of the METR RCT that quotes the study's finding that over half of AI suggestions were unusable
b5 AI Coding Tools Slowed Down Top Developers, METR Study Finds DROIDS (Substack) 2025-07 Substack newsletter coverage summarising the METR RCT's design and caveats, distinguishing greenfield/junior-developer scenarios from the tested case
b6 Vibe engineering Simon Willison's Weblog 2025-10 Simon Willison's influential independent framing of disciplined AI-assisted development as an amplifier of existing engineering practice, echoing DORA's amplifier thesis from hands-on practice rather than survey data
b7 How to Misuse & Abuse DORA Metrics GetDX Newsletter (Substack) 2022-07 Abi Noda's newsletter analysis of practitioner Bryan Finster's paper on DORA metrics becoming vanity radiators rather than improvement levers, a durable independent critique predating the AI wave
b8 Introducing the DX Core 4 GetDX Newsletter (Substack) 2024-12 Founder account of why DORA, SPACE and DevEx were combined into a fourth framework, directly addressing DORA's scope limitation to the commit-to-release window
b9 DORA's latest research on AI impact GetDX Newsletter (Substack) 2025-05 Practitioner newsletter interview with DORA's Nathen Harvey and Derek DeBellis unpacking the 2025 AI-impact findings shortly after release
b10 METR's study on how AI affects developer productivity GetDX Newsletter (Substack) 2025-07 Independent newsletter summary of the METR RCT methodology and the five factors METR identified as contributing to the slowdown
b11 DORA, SPACE, DevEx, DX Core 4 Jens Oliver Meiert (personal blog) 2025-02 Independent personal-blog critique noting that DevEx and DX Core 4 still lack the adoption and longitudinal data that give DORA and older frameworks their credibility
b12 DORA Metrics: What the Research Actually Says zbmowrey.com (personal blog) 2025 Independent practitioner blog compiling year-by-year DORA numbers (elite multipliers, AI adoption correlations, documentation multiplier) with explicit sourcing discipline
b13 Your DORA Metrics Are Lying to You: The AI ROI Crisis Tearing Through Engineering Teams Full Stack AI Engineer (Substack) 2026-03 Substack newsletter argument that AI adoption breaks the interpretability of deployment frequency, lead time and change failure rate individually, requiring supplementary signals
b14 DORA Metrics in the AI Era: Still Enough? Unblocked (company blog) 2026-06 Blog analysis contrasting the 2024 and 2025 DORA AI-adoption correlations and arguing the four keys need supplementing with a rework/churn signal and a trust/perception check
b15 DORA AI Capabilities Model: AI makes good teams better doubleSlash Blog 2026-03 Independent engineering consultancy blog dissecting the seven AI capabilities and cross-referencing them against a companion piece on why sprints lose relevance under fast AI-assisted delivery
b16 DORA the Explorer - how to unpack AI capabilities paradoxes Diginomica 2026-02 Independent enterprise-tech analyst commentary situating the 2025 AI Capabilities Model within DORA's broader research arc since 2023
b17 From Vibe Engineering to Continuous AI Continue Blog 2025-10 Independent developer-tools blog extending Simon Willison's vibe engineering framing into an organisational practice model, illustrating how the term propagated through practitioner discourse
b18 Mastering DORA Metrics in DevOps: A Guide to Optimizing Software Delivery neverblink.ai May 26, 2026 Retrieved by this lane's web search.
b19 AI Productivity Paradox: Real Developer ROI in 2025 digitalapplied.com December 26, 2025 Retrieved by this lane's web search.
b20 Research Update: Algorithmic vs. Holistic Evaluation - METR metr.org August 13, 2025 Retrieved by this lane's web search.
b21 Is AI Actually Making You Dumber and Slower? ainewstoday.substack.com Retrieved by this lane's web search.
b22 Agentic Engineering Patterns - Simon Willison's Newsletter simonw.substack.com February 27, 2026 Retrieved by this lane's web search.
b23 High Leverage | Ep. #9, The AI Coding Paradigm Shift with Simon Willison | Heavybit heavybit.com May 5, 2026 Retrieved by this lane's web search.
b24 Simon Willison on delivering AI generated code | Stefan Judis Web Development stefanjudis.com December 19, 2025 Retrieved by this lane's web search.
b25 Simon Willison’s Weblog simonwillison.net August 2, 2026 Retrieved by this lane's web search.

Financial Press

ID Title Outlet Date Significance
f1 AI slows down some experienced software developers, study finds Reuters (via Investing.com) 2025-07 Reuters wire report on the METR randomised controlled trial, the most rigorous causal evidence yet that AI coding assistance can increase task completion time despite near-universal developer belief in speedups.
f2 We are Changing our Developer Productivity Experiment Design METR 2026-02 METR's own follow-up disclosing that its second-wave experiment produced an unreliable signal due to selection effects, a methodological caveat rarely reported alongside the original 19 percent headline.
f3 Announcing the 2025 DORA Report Google Cloud Blog 2025-09 Google Cloud's official summary of the 2025 report, based on nearly 5,000 respondents and 100+ hours of interviews, introducing the AI Capabilities Model and the amplifier framing.
f4 How are developers using AI? Inside Google's 2025 DORA report Google Blog 2025-09 Google's own framing that AI adoption alone is insufficient and that a blend of technical and cultural capabilities determines outcomes, central to the accelerator vs amplifier distinction.
f5 DORA | DORA 2025: Year in review DORA 2025 DORA's own recap distinguishing the September State of AI-assisted Software Development report from the December AI Capabilities Model companion, useful for tracking methodology and publication cadence.
f6 Sources: Cursor in talks to raise $2B+ at $50B valuation as enterprise growth surges TechCrunch (citing Bloomberg) 2026-04 TechCrunch report citing Bloomberg on Cursor/Anysphere's revenue trajectory (over 2 billion in annualised revenue by February 2026, forecasting over 6 billion by end-2026), illustrating the scale of capital betting on AI-coding productivity claims.
f7 Cursor's Anysphere nabs $9.9B valuation, soars past $500M ARR TechCrunch (citing Bloomberg) 2025-06 TechCrunch report citing Bloomberg's original reporting on Anysphere's 900 million dollar raise, documenting investor conviction in AI-coding tools well ahead of independent productivity verification.
f8 Anysphere, maker of AI coding tool Cursor, confirms $900M raise led by Thrive Capital Crunchbase News 2025-06 Crunchbase News piece corroborating Bloomberg's description of Anysphere as "the fastest growing startup ever", tying the AI-coding capital story to specific investor names.
f9 AI in Software Development: Productivity at the Cost of Code Quality? DevOps.com (citing Wall Street Journal) 2025-12 Trade coverage quoting the Wall Street Journal's interview with MIT's Armando Solar-Lezama comparing AI coding assistance to a credit card that lets teams "accumulate technical debt", linking financial-press framing to DevOps quality concerns.
f10 How AI is already shaking up the job market Morning Brew 2025-05 Reports TD Bank's comments to the Wall Street Journal that it now screens for abstract coding and prompt-engineering ability, and Salesforce's claim of a 30% productivity boost behind its engineering hiring freeze.
f11 AI's productivity paradox AOL (Business Insider) 2026-06 Cites Moody's chief economist Mark Zandi arguing AI's productivity lift will emerge gradually rather than immediately, framing the gap between capital markets' AI enthusiasm and measured output.
f12 The Economics of Intelligence Citadel Securities 2026-05 Citadel Securities analysis applying the Jevons Paradox to software engineering, arguing cheaper AI-assisted coding could expand rather than shrink developer demand, a market-structure argument financial analysts use to explain flat headcount despite productivity claims.
f13 Research: Quantifying GitHub Copilot's impact in the enterprise with Accenture GitHub Blog 2024-05 Primary vendor-commissioned study (GitHub/Accenture) on enterprise Copilot telemetry, useful as a vendor claim requiring independent corroboration against DORA and METR findings.
f14 AI is eroding code quality states new in-depth report DevClass 2025-02 devclass coverage of GitClear's findings that duplicated code blocks rose eightfold in 2024, providing an independent technology-press read on the code-quality telemetry debate.
f15 RDEL #112: What's AI's impact on software delivery performance? (2025 DORA Report) Research-Driven Engineering Leadership (Substack) 2025 Practitioner analysis quantifying DORA's Bayesian modelling approach (nearly 5,000 professionals, 100+ hours of interviews) and its finding that AI adoption improved throughput but continued to increase instability in 2025.
f16 How Generative AI Is Changing Software Development: Key Insights from the DORA Report opslevel.com Retrieved by this lane's web search.
f17 Top takeaways from the 2024 DORA Report – sponsored by Gearset! | DevOps Launchpad devopslaunchpad.com October 23, 2024 Retrieved by this lane's web search.
f18 DORA 2024: AI Is Speeding Up Delivery and Hurting Stability at the Same Time | Ali Safari alisafari.space May 14, 2026 Retrieved by this lane's web search.
f19 METR on X: "We ran a randomized controlled trial to see how much AI coding tools speed up experienced open-source developers. The results surprised us: Developers thought they were 20% faster with AI tools, but they were actually 19% slower when they had access to AI than when they didn't." / X x.com July 10, 2025 Retrieved by this lane's web search.
f20 My Participation in the METR AI Productivity Study | Domenic Denicola domenic.me July 15, 2025 Retrieved by this lane's web search.
f21 IACDM: Interactive Adversarial Convergence Development Methodology -- A Structured Framework for AI-Assisted Software Development arxiv.org Retrieved by this lane's web search.
f22 A METR Study Reveals that AI Slows Down Experienced Developers - ActuIA actuia.com July 16, 2025 Retrieved by this lane's web search.
f23 AI Coding Statistics - Adoption, Productivity & Market Metrics getpanto.ai August 7, 2026 Retrieved by this lane's web search.

Frontier Lab & Model News

ID Title Outlet Date Significance
t1 Details about METR's evaluation of OpenAI GPT-5 METR 2025-08 Primary METR pre-deployment evaluation report giving GPT-5's autonomous coding time-horizon figures used to benchmark frontier capability growth against DORA-era instability findings.
t2 METR's GPT-4.5 pre-deployment evaluations METR 2025-02 Earlier METR evaluation establishing the time-horizon methodology later applied to GPT-5 and subsequent models, showing the capability trendline behind AI-coding adoption.
t3 METR's Evaluation of GPT-5 Alignment Forum 2025-08 Cross-posted METR researcher commentary giving additional statistical detail (bootstrap comparisons, confidence intervals) on the GPT-5 autonomy evaluation.
t4 OpenAI GPT-5 System Card OpenAI (arXiv mirror) 2025 Official system card documenting METR's external evaluation conclusions on GPT-5-thinking's autonomy and AI R&D speed-up risk, the primary-source counterpart to METR's own blog report.
t5 GPT-5.6 Preview System Card - External Evaluations for AI Self-Improvement, METR OpenAI Deployment Safety Hub 2026 OpenAI's Deployment Safety Hub documentation of continued METR involvement in evaluating successive GPT-5.x releases for autonomy and destructive-action risks.
t6 System Card: Claude Opus 5 Anthropic 2026-07 Anthropic's system card giving its own AI R&D capability index (ECI) comparison for Claude Opus 5, the frontier-lab counterpart to METR's independent time-horizon tracking.
t7 2025 DORA State of AI Assisted Software Development Report intelligence.theregister.com Retrieved by this lane's web search.
t8 AI Dev: The 2024 DORA Report Reviewed | by Julian B | Medium medium.com November 27, 2024 Retrieved by this lane's web search.
t9 DevOps in 2024: Essential Insights from the DORA Report | Upsun upsun.com November 19, 2025 Retrieved by this lane's web search.
t10 AI Writes 41% of Code: The Churn and Tech-Debt Data particula.tech June 5, 2026 Retrieved by this lane's web search.
t11 GPT-5.3 and Claude Opus 4.6: More System Card Shenanigans ignorance.ai February 11, 2026 Retrieved by this lane's web search.
t12 Understanding The 4 DORA Metrics And Top Findings From 2024/25 DORA Report | Octopus Deploy octopus.com Retrieved by this lane's web search.
t13 What are the DORA metrics? dynatrace.com October 17, 2025 Retrieved by this lane's web search.
t14 Understanding DORA Metrics for DevOps Performance oneuptime.com February 20, 2026 Retrieved by this lane's web search.

VC & Analyst Reports

ID Title Outlet Date Significance
v1 Gartner Magic Quadrant for AI Code Assistants Gartner 2025-09 Gartner's primary 2025 vendor evaluation of 14 AI code-assistant providers, the analyst-radar reference point for this market.
v2 The AI software development market map CB Insights 2025-10 CB Insights market map of code-generation, review and testing vendors, with ARR growth figures for Anysphere/Cursor, Replit and Lovable.
v3 The AI agent market map CB Insights 2026-03 CB Insights' broader agent market map identifying software development as one of the two most-traction AI agent categories.
v4 Developer Tools 2.0 Sequoia Capital 2023-05 Sequoia's foundational 2023 investment thesis on Copilot-style AI coding tools, illustrating the early accelerator framing later complicated by DORA/METR data.
v5 Services: The New Software Sequoia Capital 2026-03 Sequoia's 2026 thesis pivot toward AI 'autopilots' replacing labour budgets rather than just tooling, showing the firm's thesis evolution beyond simple coding-assistant acceleration.
v6 Unleash developer productivity with generative AI McKinsey & Company 2023-06 McKinsey's original 2023 lab study of 40+ developers finding productivity and flow gains, a precursor to later organisation-level scepticism.
v7 AI in software development: boosting productivity McKinsey & Company 2025-12 McKinsey's 2025-2026 interview citing Jellyfish vendor data on adoption thresholds and productivity gains, illustrating vendor-sourced numbers entering consultancy commentary.
v8 From Pilots to Payoff: Generative AI in Software Development Bain & Company 2025-09 Bain's primary 2025 Technology Report finding only 10-15% productivity gains from AI coding assistants and arguing for lifecycle-wide redesign.
v9 Technology Report 2025 - Technology Industry Trends Bain & Company 2025-09 Bain's topic hub for its sixth annual Technology Report, framing AI productivity gains against required process change.
v10 AI coding hype overblown, Bain shrugs The Register 2025-09 Independent press summary of Bain's 2025 findings on low developer adoption and modest 10-15% productivity gains.
v11 Predictions 2025: GenAI Reality Bites Back For Software Developers Forrester 2024-10 Forrester's 2025 predictions, including that developers spend only ~24% of time coding and a prediction of a failed 50%-developer-replacement attempt.
v12 What AI-Enhanced Software Development Means For Technology Executives Forrester 2025-09 Forrester's 2025 developer-survey commentary on uneven AI-assistant adoption across SDLC stages ('TuringBots').
v13 AI Is Amplifying Software Engineering Performance, Says the 2025 DORA Report InfoQ 2026-03 Independent trade-press synthesis of the 2025 DORA report's throughput-up/stability-down finding and the amplifier framing.
v14 Enterprise Tech Investments & Team Overview Andreessen Horowitz (a16z) 2025 a16z's enterprise AI investment hub, tracking enterprise gen-AI spend growth and top AI-native productivity companies relevant to market sizing.
v15 Productivity and Pitfalls in AI Coding DORA 2025 | by Tamanna | Medium medium.com September 29, 2025 Retrieved by this lane's web search.
v16 AI Won’t Fix Broken Systems: Lessons from the 2025 DORA Report - Aviator Blog aviator.co June 7, 2026 Retrieved by this lane's web search.
v17 DORA Metrics Engineering Effectiveness: AI Impact in 2026 blog.exceeds.ai July 30, 2026 Retrieved by this lane's web search.
v18 DORA Report 2025: How AI Adoption Shapes DevOps and Software Teams - Opsera opsera.ai November 5, 2025 Retrieved by this lane's web search.
v19 2025 DORA State of AI-assisted Software Development Report cloud.google.com Retrieved by this lane's web search.
v20 McKinsey is Still Talking about Eng Prod - Here's Why faros.ai February 26, 2026 Retrieved by this lane's web search.
v21 The Productivity Paradox: Why AI Developers Feel Faster But Deliver Slower | by Eran Swears | Medium medium.com December 9, 2025 Retrieved by this lane's web search.
v22 McKinsey State of AI 2025: What It Means for Engineering Leaders colabsoftware.com April 10, 2026 Retrieved by this lane's web search.
v23 McKinsey & Company: Measuring AI Developer Productivity & Attrition Risk 2026 codeninety.com February 18, 2026 Retrieved by this lane's web search.

We use analytics cookies to understand site usage and improve the service. We do not use marketing cookies.