Research · Blogs & Independent Thinkers
Back to sweepResearch sweep · deep · 2023 – 2026
DORA Metrics in the AI Era
How the DORA four keys (deployment frequency, lead time for changes, change failure rate, failed deployment recovery time) plus the 2024 rework-rate metric are measured and improved in practice, and how AI-assisted development has shifted them, September 2023 to September 2026: DORA State of DevOps 2023 and 2024, the 2025 State of AI-assisted Software Development and its AI Capabilities Model, the METR developer RCT, GitClear code-churn data, and the SPACE and DevEx frameworks
- Claude Fable 5
- tech
- academic
- blogs
- financial
- frontier
- vc
Synthesised 2026-09-09
Narrative
Independent commentary on the DORA metrics largely treats the annual report itself as a text to be parsed rather than a settled scoreboard. RedMonk analyst Rachel Stephens' close reading of the 2024 Accelerate State of DevOps report is the most cited independent analysis: she flags that DORA reclassified time to restore as a throughput measure and added rework rate to stability, and floats a theory-of-constraints reading that "writing code isn't the bottleneck to deploying reliable applications." Her 2025 follow-up tracks the reversal that made the throughput sign flip positive while noting DORA's own language that AI adoption "not only fails to fix instability, it is currently associated with increasing instability," and records that in 2025 usage rose to 90% of respondents at a median two hours per day.
On the METR randomised controlled trial, independent bloggers treat the 16-developer, 246-task sample with more care than headline aggregators. A LessWrong analysis (cross-posted from the "Don't Worry About the Vase" circle) argues the slowdown is real but may not generalise beyond experienced developers working in codebases they already know deeply, citing Emmett Shear's point that general LLM experience does not transfer cleanly to tool-specific fluency with Cursor. The skeptical outlet Pivot to AI frames the result bluntly, noting developers predicted a 24% speedup, felt 20% faster, and were measured as 19% slower, and highlights that "over half the AI suggestions were not usable." METR's own Substack update in February 2026 is notable for methodological candour: a follow-up study begun in August 2025 was abandoned as an unreliable signal because too many developers refused to be randomised into the no-AI arm, which the team says "likely biases downwards our estimate of AI-assisted speedup."
Practitioner newsletters occupy the space between DORA and vendor dashboards. Abi Noda and Laura Tacho's GetDX newsletter has tracked DORA methodology changes since before the AI era, including a 2022 piece on how organisations misuse DORA as a vanity metric rather than an improvement lever, and more recently argues DORA's four keys alone under-measure AI-assisted teams because deployment frequency inflates with AI-generated scaffolding while review overhead for AI code often exceeds that for human-written code. Gergely Orosz's Pragmatic Engineer newsletter carried the original DevEx framework interview with Nicole Forsgren, Margaret-Anne Storey, Michaela Greiler and Abi Noda, which explicitly frames DevEx as filling the gap "before and after" the commit-to-release window DORA measures. Independent engineering blogs (Jens Oliver Meiert, personal blog) are candid that DevEx and DX Core 4 remain thin on real-world adoption data compared with DORA's decade of longitudinal survey history.
GitClear's code-churn findings are widely relayed by independent technical bloggers (Rob Bowley, jonas.rs) as a corroborating signal for DORA's instability findings, though all treat it as vendor telemetry from a single company's own review-tool customer base rather than an independent academic result. Simon Willison's "vibe engineering" post, arguably the most-cited independent framing of 2025 outside DORA itself, arrives at a similar amplifier conclusion through hands-on practice rather than survey data: AI tools multiply the effect of existing engineering discipline (testing, planning, code review, documentation) rather than substituting for it, a framing several practitioner blogs explicitly map onto DORA's "AI is an amplifier" language.
Sources
| ID | Title | Outlet | Date | Significance |
|---|---|---|---|---|
| b1 | DORA 2025: Measuring Software Delivery After AI | RedMonk | 2025-12 | Follow-up independent analysis tracking the 2025 reversal in AI-throughput correlation and continued instability correlation, with direct quotes from the DORA report |
| b2 | On METR's AI Coding RCT | LessWrong | 2025-07 | LessWrong community analysis testing the METR RCT's generalisability, weighing the experience-with-tool hypothesis against the raw slowdown finding |
| b3 | We are Changing our Developer Productivity Experiment Design | METR (Substack) | 2026-02 | METR's own Substack admission that a 2026 follow-up study produced an unreliable signal due to participation bias, an important caveat on replication |
| b4 | AI coders think they're 20% faster - but they're actually 19% slower | Pivot to AI | 2025-07 | Skeptical independent blog treatment of the METR RCT that quotes the study's finding that over half of AI suggestions were unusable |
| b5 | AI Coding Tools Slowed Down Top Developers, METR Study Finds | DROIDS (Substack) | 2025-07 | Substack newsletter coverage summarising the METR RCT's design and caveats, distinguishing greenfield/junior-developer scenarios from the tested case |
| b6 | Vibe engineering | Simon Willison's Weblog | 2025-10 | Simon Willison's influential independent framing of disciplined AI-assisted development as an amplifier of existing engineering practice, echoing DORA's amplifier thesis from hands-on practice rather than survey data |
| b7 | How to Misuse & Abuse DORA Metrics | GetDX Newsletter (Substack) | 2022-07 | Abi Noda's newsletter analysis of practitioner Bryan Finster's paper on DORA metrics becoming vanity radiators rather than improvement levers, a durable independent critique predating the AI wave |
| b8 | Introducing the DX Core 4 | GetDX Newsletter (Substack) | 2024-12 | Founder account of why DORA, SPACE and DevEx were combined into a fourth framework, directly addressing DORA's scope limitation to the commit-to-release window |
| b9 | DORA's latest research on AI impact | GetDX Newsletter (Substack) | 2025-05 | Practitioner newsletter interview with DORA's Nathen Harvey and Derek DeBellis unpacking the 2025 AI-impact findings shortly after release |
| b10 | METR's study on how AI affects developer productivity | GetDX Newsletter (Substack) | 2025-07 | Independent newsletter summary of the METR RCT methodology and the five factors METR identified as contributing to the slowdown |
| b11 | DORA, SPACE, DevEx, DX Core 4 | Jens Oliver Meiert (personal blog) | 2025-02 | Independent personal-blog critique noting that DevEx and DX Core 4 still lack the adoption and longitudinal data that give DORA and older frameworks their credibility |
| b12 | DORA Metrics: What the Research Actually Says | zbmowrey.com (personal blog) | 2025 | Independent practitioner blog compiling year-by-year DORA numbers (elite multipliers, AI adoption correlations, documentation multiplier) with explicit sourcing discipline |
| b13 | Your DORA Metrics Are Lying to You: The AI ROI Crisis Tearing Through Engineering Teams | Full Stack AI Engineer (Substack) | 2026-03 | Substack newsletter argument that AI adoption breaks the interpretability of deployment frequency, lead time and change failure rate individually, requiring supplementary signals |
| b14 | DORA Metrics in the AI Era: Still Enough? | Unblocked (company blog) | 2026-06 | Blog analysis contrasting the 2024 and 2025 DORA AI-adoption correlations and arguing the four keys need supplementing with a rework/churn signal and a trust/perception check |
| b15 | DORA AI Capabilities Model: AI makes good teams better | doubleSlash Blog | 2026-03 | Independent engineering consultancy blog dissecting the seven AI capabilities and cross-referencing them against a companion piece on why sprints lose relevance under fast AI-assisted delivery |
| b16 | DORA the Explorer - how to unpack AI capabilities paradoxes | Diginomica | 2026-02 | Independent enterprise-tech analyst commentary situating the 2025 AI Capabilities Model within DORA's broader research arc since 2023 |
| b17 | From Vibe Engineering to Continuous AI | Continue Blog | 2025-10 | Independent developer-tools blog extending Simon Willison's vibe engineering framing into an organisational practice model, illustrating how the term propagated through practitioner discourse |
| b18 | Mastering DORA Metrics in DevOps: A Guide to Optimizing Software Delivery | neverblink.ai | May 26, 2026 | Retrieved by this lane's web search. |
| b19 | AI Productivity Paradox: Real Developer ROI in 2025 | digitalapplied.com | December 26, 2025 | Retrieved by this lane's web search. |
| b20 | Research Update: Algorithmic vs. Holistic Evaluation - METR | metr.org | August 13, 2025 | Retrieved by this lane's web search. |
| b21 | Is AI Actually Making You Dumber and Slower? | ainewstoday.substack.com | Retrieved by this lane's web search. | |
| b22 | Agentic Engineering Patterns - Simon Willison's Newsletter | simonw.substack.com | February 27, 2026 | Retrieved by this lane's web search. |
| b23 | High Leverage | Ep. #9, The AI Coding Paradigm Shift with Simon Willison | Heavybit | heavybit.com | May 5, 2026 | Retrieved by this lane's web search. |
| b24 | Simon Willison on delivering AI generated code | Stefan Judis Web Development | stefanjudis.com | December 19, 2025 | Retrieved by this lane's web search. |
| b25 | Simon Willison’s Weblog | simonwillison.net | August 2, 2026 | Retrieved by this lane's web search. |