Research · DORA Metrics in the AI Era
Back to researchResearch sweep · deep · 2023 – 2026
DORA Metrics in the AI Era
How the DORA four keys (deployment frequency, lead time for changes, change failure rate, failed deployment recovery time) plus the 2024 rework-rate metric are measured and improved in practice, and how AI-assisted development has shifted them, September 2023 to September 2026: DORA State of DevOps 2023 and 2024, the 2025 State of AI-assisted Software Development and its AI Capabilities Model, the METR developer RCT, GitClear code-churn data, and the SPACE and DevEx frameworks
Synthesised 2026-09-09
DORA in the AI era: the metrics held, the practices got louder
Overview
The four keys survived AI. Everything around them shifted. Between September 2023 and September 2026 DORA renamed one metric (time to restore became failed deployment recovery time, redefined in 2023 to isolate failures caused by a production change), reclassified another (recovery time moved from stability to throughput in 2024), added a fifth (rework rate, unplanned deployments fixing user-facing bugs), and in 2025 rebranded the whole programme as the State of AI-assisted Software Development, retiring low-to-elite performance tiers in favour of seven team archetypes. Sources: DORA / Google Cloud (2026) (↗); Continuous Delivery Foundation (2025) (↗); redmonk.com (2024) (↗); RedMonk (2025) (↗)
The defining shift of the past eighteen months is the collapse of the simple speed-up narrative. DORA 2024 measured a 25% increase in AI adoption as correlating with a 7.2% reduction in delivery stability and a 1.5% reduction in throughput, even as individual productivity and flow rose. METR's randomised trial then found experienced developers 19% slower with AI while believing themselves 20% faster. By 2025 DORA had replaced trade-off language with an amplifier thesis: AI magnifies the strengths of well-run organisations and the dysfunctions of struggling ones. Sources: dora.dev (n.d.) (↗); metr.org (2025) (↗); Google Cloud Blog (2025) (↗)
For a leader deciding what to change on Monday, the practical message is conservative. The capabilities that predicted the four keys before AI (small batches, version control discipline, continuous testing, platform quality) are the same ones DORA's 2025 AI Capabilities Model found determine whether AI helps or harms delivery. AI raises the price of skipping them, because it makes it trivially easy to enlarge batch size, duplicate code, and outrun review capacity. Sources: Google Research (2025) (↗); Substack (practitioner) (2024) (↗)
Timeline
- DORA redefines recovery metric as failed deployment recovery time, isolating change-caused failures
- GitHub and Accenture publish enterprise Copilot ROI figures, setting the vendor benchmark for gains
- DORA 2024 measures AI adoption reducing throughput and stability
- Rework rate added as fifth metric, recovery time reclassified as throughput
- GitClear reports copy-pasted code exceeding refactored code for the first time
- Cursor maker Anysphere valued at 9.9 billion dollars, capital fully committed to the acceleration thesis
- METR RCT finds 19% slowdown against a forecast 24% speedup
- DORA rebrands as State of AI-assisted Software Development, ships the seven-capability AI model and seven team archetypes
- Gartner publishes first Magic Quadrant for AI code assistants
- METR abandons its follow-up study after developers refuse the no-AI arm
- GitClear extends telemetry, block duplication up 81% on the 2023 baseline
- Vibe coding gives way to disciplined context engineering per Thoughtworks
Key Findings
Throughput and stability are now framed as companions, not a trade-off, and the metric groupings changed to enforce that. DORA 2024 formally paired deployment frequency and lead time as throughput and change failure rate and recovery time as stability, then moved recovery time into throughput and treated change failure rate as a proxy for rework, adding deployment rework rate to measure it directly. RedMonk's Rachel Stephens flagged this as the report's most consequential structural move. The decade-long finding behind it is that high performers achieve both at once, so any local improvement that trades one for the other is a mis-optimisation. Sources: redmonk.com (2024) (↗); Continuous Delivery Foundation (2025) (↗); DORA / Google Cloud (2026) (↗)
The strongest levers per metric are boring and repeated across report years. Lead time responds to batch size and trunk-based version control practice: DORA's explanation for the 2024 AI penalty was precisely that AI enlarges changesets, and "DORA's data has consistently shown that larger changesets introduce risk". Change failure rate and rework respond to continuous testing and review discipline. Recovery time responds to version control hygiene and rollback capability, which is why "strong version control practices" appears verbatim in the 2025 AI Capabilities Model. The Monday action is mechanical: cap PR size, keep AI-generated work on short-lived branches, and refuse to let scaffolding volume masquerade as throughput. Sources: thenewstack.io (2024) (↗); Google Research (2025) (↗); aviator.co (2026) (↗)
DORA 2024's headline was a measured contradiction, and it replicated only partially. With roughly 39,000 respondents and about 3,000 answering AI questions, the 2024 Bayesian model tied a 25% AI adoption increase to a 7.2% stability drop, a 1.5% throughput drop, and a 2.1% individual productivity gain. The 2025 wave (near 5,000 respondents, over 100 hours of interviews) flipped the throughput sign positive while instability persisted; DORA's own language was that AI "not only fails to fix instability, it is currently associated with increasing instability". The stability finding has now held across two survey waves. The throughput finding has not, so treat it as unsettled. Sources: dora.dev (n.d.) (↗); RedMonk (2025) (↗); Research-Driven Engineering Leadership (Substack) (2025) (↗)
METR's RCT is the field's only controlled evidence, and it broke the self-report link. Sixteen experienced open-source developers, 246 real tasks in repositories they had maintained for an average of five years, mostly Cursor Pro with Claude 3.5/3.7 Sonnet: measured completion time rose 19% (95% CI roughly -40% to -2%) against a forecast 24% speedup and a post-hoc estimate of 20% faster. Independent readings from LessWrong and others argue the result may not generalise beyond experts in familiar code, and METR itself now labels it historical, stating in February 2026 that it believes developers are more sped up by early-2026 AI than the study measured. The durable lesson is narrower and more useful: developer self-report is an unreliable instrument for AI productivity, so measure delivery outcomes instead. Sources: arXiv / METR (2025) (↗); LessWrong (2025) (↗); METR (Substack) (2026) (↗); Reuters (via Investing.com) (2025) (↗)
GitClear's telemetry corroborates the instability side, not the productivity side. Across 211 million changed lines (2020 to 2024), copy-pasted code exceeded moved (refactored) code for the first time in 2024, duplicated blocks rose roughly eightfold, and two-week churn climbed from about 3.1% in 2020 to 5.7% in 2024, with moved code falling from about 25% of changed lines in 2021 to under 10%. The 2026 follow-up found block duplication up 81% on the 2023 baseline, and, tellingly, that heavy AI users' 4 to 10x output advantage over non-users mostly predates their AI adoption, with a more modest 25% velocity gain against their own past output. This is vendor telemetry from one company's customer base, but it is the only commit-level dataset at this scale and it points the same direction as DORA's stability finding. Sources: DevClass (2025) (↗); particula.tech (2026) (↗); DevOps.com (citing Wall Street Journal) (2025) (↗)
The 2025 AI Capabilities Model is mostly the old capability list with an AI prefix, and that is its point. Built from 78 interviews and validated against the near-5,000 sample, its seven capabilities are a clear and communicated AI stance, healthy data ecosystems, AI-accessible internal data, strong version control practices, working in small batches, user-centric focus, and quality internal platforms, selected from fifteen candidates for showing a statistically significant interaction with outcomes. Three of the seven (small batches, version control, platforms) are pre-existing DORA capabilities. The genuinely new ones concern making internal context available to AI tools and telling people what they are allowed to do with them, both cheap to implement and, per DORA, both currently rare. Sources: Google Research (2025) (↗); services.google.com (n.d.) (↗); Google Cloud Blog (2025) (↗)
Platform engineering shows the same bifurcation as AI itself. DORA 2024 found internal platforms lifted individual productivity by roughly 8% and organisational performance by roughly 6% while reducing throughput by about 8% and stability by about 14%. The pattern (individual gains, delivery losses) mirrors the AI finding, suggesting a general failure mode: tools that make individuals feel faster while delivery discipline erodes. A platform is a lever only when paired with the batch-size and testing practices it is supposed to encode. Sources: thenewstack.io (2024) (↗); redmonk.com (2024) (↗)
SPACE, DevEx and DX Core 4 now bracket DORA rather than compete with it. The DevEx framework (Noda, Storey, Forsgren, Greiler, ACM Queue 2023) explicitly measures the experience "before and after" the commit-to-production window DORA covers, via feedback loops, cognitive load and flow. DX Core 4 (2024) consolidates the lot. GetDX's practitioner argument is that the four keys under-measure AI-assisted teams because deployment frequency inflates with AI scaffolding while review overhead for AI code often exceeds that for human code, so review load and cognitive load belong on the dashboard next to the five keys. Sources: The Pragmatic Engineer (Gergely Orosz) (2024) (↗); GetDX Newsletter (Substack) (2024) (↗); Unblocked (company blog) (2026) (↗); Jens Oliver Meiert (personal blog) (2025) (↗)
A good team in 2026 looks like DORA's "harmonious high-achievers" archetype, and its numbers are behavioural, not threshold-based. Ninety per cent of respondents use AI daily at a median two hours, yet 30% report little or no trust in AI output, and the good teams treat that distrust as process: small batches enforced on AI-generated changes, rework rate tracked as a first-class stability signal, internal data wired into tools, and an explicit organisational AI stance. Thoughtworks' Rachel Laycock reports vibe coding "practically disappeared" by late 2025 in favour of context engineering, which is the same conclusion Simon Willison reached from practice: AI multiplies existing discipline in testing, planning, review and documentation. Elite multipliers such as 127x lead time still circulate from the 2024 report, but the tier structure behind them no longer exists, so quote them with a year attached or not at all. Sources: RedMonk (2025) (↗); cloud.google.com (n.d.) (↗); Thoughtworks (2026) (↗); Simon Willison's Weblog (2025) (↗)
Evidence & Data
The load-bearing numbers, by source type. DORA survey (organisational, correlational): 2024, 25% more AI adoption maps to -7.2% stability, -1.5% throughput, +2.1% individual productivity, on ~39,000 respondents; 2025, throughput turns positive, instability persists, 90% daily AI use, 30% low trust, on ~5,000 respondents. Controlled trial (individual, causal): METR, 16 developers, 246 tasks, +19% completion time against forecast -24%. Commit telemetry (codebase, observational): GitClear, churn 3.1% to 5.7% (2020 to 2024), duplication up roughly eightfold in 2024, up 81% on 2023 by 2026, cloned code linked to 15 to 50% more defects. Sources: dora.dev (n.d.) (↗); Google Blog (2025) (↗); metr.org (2025) (↗); particula.tech (2026) (↗)
These are different units of analysis and should never be merged. Cui and colleagues' pooled field experiments across Microsoft, Accenture and a Fortune 100 firm (4,867 developers) found a 26.08% increase in completed tasks, with larger gains for juniors, while Peng's original Copilot RCT found 55.8% on a single scripted task. Both measure task output, not delivery. Bain puts realised team-level gains at 10 to 15%, noting coding is only 25 to 35% of the lifecycle; Gartner forecasts 30% "through 2028" as projection, not measurement; McKinsey's Jellyfish-partnered claim of over 110% gains at 80 to 100% adoption is vendor-sourced and unreplicated. Sources: cmustatistics.github.io (2025) (↗); Bain & Company (2025) (↗); Gartner (2025) (↗); McKinsey & Company (2025) (↗)
The market priced the bull case regardless: Anysphere at 9.9 billion dollars in June 2025 past 500 million ARR, past 2 billion ARR by early 2026, with talks at a 50 billion valuation. Sources: TechCrunch (citing Bloomberg) (2025) (↗); TechCrunch (citing Bloomberg) (2026) (↗)
Signals & Tensions
Measured versus asserted. DORA's Bayesian survey model, METR's RCT and GitClear's telemetry form the measured tier. Vendor telemetry (Faros, Jellyfish), Copilot ROI studies and Magic Quadrant placements form the asserted tier, and consistently report larger, more favourable numbers. The gap between the two tiers runs 30 to 50 percentage points. Sources: GitHub Blog (2024) (↗); faros.ai (2026) (↗); The Register (2025) (↗)
Instrumentation drift. DORA's canonical definitions come from self-report, while tool vendors instrument the same names through CI/CD data with materially different choices, such as whether rollbacks count as deployments or lead time starts at ticket rather than commit. Both errors flatter the numbers, and the survey-versus-telemetry gap is largely unquantified. Sources: getdx.com (2026) (↗); dynatrace.com (2025) (↗)
The METR refusal problem is itself a finding. The follow-up died because developers would not accept randomisation into a no-AI arm, which biases future estimates downward and means the cleanest study design in the field may be unrepeatable. Sources: METR (Substack) (2026) (↗); METR (2026) (↗)
Capability trend versus delivery trend. METR's time-horizon evaluations show autonomous coding capability climbing steadily (GPT-5 at roughly 2h15m at 50% reliability against o3's 1h30m) while team-level stability telemetry worsens with adoption. Both can be true; the bottleneck has moved to integration, review and rework. Sources: METR (2025) (↗); metr.org (2025) (↗)
Underreported: skill formation. A 2026 randomised study found AI use impaired conceptual understanding and debugging without average efficiency gains, a cost no delivery metric currently captures. Sources: ainewstoday.substack.com (n.d.) (↗)
Open Questions
Whether the 2025 throughput reversal is a real learning effect or an artefact of a changed sample and instrument, since DORA's waves are not a panel and question wording shifted with the rebrand. Sources: DORA (2025) (↗)
Whether rework rate and the AI-era proposals (AI-attributable rework, acceptance rate, trust, review load) survive a second survey wave, or join the churn of retired constructs like the elite tier. Sources: RedMonk (2025) (↗); GetDX Newsletter (Substack) (2025) (↗)
Whether GitClear's duplication and churn trends replicate outside its own customer base; the first independent RCTs on AI-assisted maintainability are only now appearing on arXiv. Sources: DevClass (2025) (↗)
How much of the individual-versus-organisational gap is review bottleneck, measurable as review load, versus genuine quality decay, measurable as defects, since the interventions differ. Sources: Unblocked (company blog) (2026) (↗)
Whether a stable elite profile for AI-assisted teams exists at all, given the archetypes are one wave old and the cluster analysis that once defined elite has already failed once, in 2022. Sources: cloud.google.com (2020) (↗); Octopus Deploy blog (2024) (↗)
Whether any credible RCT of AI-era delivery is still possible now that developers refuse control conditions, or whether the field is permanently limited to observational telemetry and survey correlation. Sources: METR (Substack) (2026) (↗)
The Monday takeaway holds either way: the five keys still measure the right things, AI has not changed which practices move them, and the teams winning in 2026 are the ones that got stricter about small batches and review exactly when their tools made it easiest to abandon both.
Sources
Summary: ↑ Back to summary
Tech Industry & Practitioner
| ID | Title | Outlet | Date | Significance |
|---|---|---|---|---|
| p1 | DORA: AI boosting productivity, hindering delivery | Substack (practitioner) | 2024 | Highlights the counterintuitive 2024 finding that AI improved code quality/review speed indicators while lowering lead time, throughput and change failure rate. |
| p2 | Introducing the DORA AI Capabilities Model: 7 keys to succeeding in AI-assisted software development | Google Research | 2025-09 | Google Research write-up describing the three-phase methodology (78 interviews, validated survey, ~5,000-respondent analysis) behind the capabilities model. |
| p3 | From adoption to impact: Putting the DORA AI Capabilities Model to work | Google Cloud Blog | 2025-12 | DORA lead Nathen Harvey's own framing of AI as an amplifier, citing 90% daily AI adoption and cluster analysis showing outcome variation across teams. |
| p4 | The 2024 DevOps Performance Clusters | Octopus Deploy blog | 2024-11 | Explains how DORA's elite/high/medium/low clusters are derived statistically each year and cautions against treating any single year's thresholds as fixed benchmarks. |
| p5 | The DORA 4 key metrics become 5 | Continuous Delivery Foundation | 2025-10 | CD Foundation explainer detailing exactly how rework rate is calculated and why recovery time moved from stability to throughput. |
| p6 | DORA | A history of DORA's software delivery metrics | DORA / Google Cloud | 2026-01 | DORA's own chronology of metric changes since 2014, including the 2018 addition of availability and the 2024 addition of deployment rework rate. |
| p7 | DORA metrics | Technology Radar | Thoughtworks Technology Radar | 2026-04 | Thoughtworks' practitioner take arguing that in the AI era, lines-of-code output is a misleading productivity signal versus delivery flow and stability. |
| p8 | Thoughtworks Technology Radar Highlights The Rapid Evolution of AI Assistance in 2025 | Thoughtworks | 2025-11 | CTO Rachel Laycock's account of AI antipatterns (shadow IT, complacency with AI code) and the shift from vibe coding to context engineering across 2025. |
| p9 | From vibe coding to context engineering: 2025 in software development | Thoughtworks | 2026-06 | Thoughtworks practitioner reflection connecting context engineering practice to reduced rework and better AI code reliability. |
| p10 | How to Measure Developer Productivity: DORA vs SPACE vs GSM | ShiftMag | 2026-03 | Situates DORA against the SPACE (2021) and DevEx (April 2023, ACM Queue) frameworks by original named authors including Nicole Forsgren and Margaret-Anne Storey. |
| p11 | A new way to measure developer productivity – from the creators of DORA and SPACE | The Pragmatic Engineer (Gergely Orosz) | 2024-05 | Interview with DevEx framework authors explaining the three-dimension (feedback loops, cognitive load, flow state) model and its lineage from DORA and SPACE. |
| p12 | DORA 2025 AI Capabilities: How to Boost AI Impact | kt.team | December 9, 2025 | Retrieved by this lane's web search. |
| p13 | 2025 AI Capabilities Model | meriroos.ee | Retrieved by this lane's web search. | |
| p14 | Quantifying the Expectation-Realisation Gap for Agentic AI Systems | arxiv.org | Retrieved by this lane's web search. | |
| p15 | DORA Metrics Explained (2026): Four Keys & Benchmarks | taskade.com | June 9, 2026 | Retrieved by this lane's web search. |
| p16 | Use Four Keys metrics like change failure rate to measure your DevOps performance | Google Cloud Blog | cloud.google.com | September 22, 2020 | Retrieved by this lane's web search. |
| p17 | DORA Metrics: Measure Open DevOps Success with 4 Key Indicators | waydev.co | December 16, 2025 | Retrieved by this lane's web search. |
| p18 | What Elite Performing Teams Measure | An Introduction to DORA Metrics | Lean TECHniques | leantechniques.com | March 2, 2026 | Retrieved by this lane's web search. |
| p19 | DORA metrics: the complete guide to measuring DevOps performance in the AI era | getdx.com | July 13, 2026 | Retrieved by this lane's web search. |
| p20 | DORA Metrics Benchmarks 2026: Where Does Your Team Stand? | Gitrecap Blog | gitrecap.com | April 9, 2026 | Retrieved by this lane's web search. |
| p21 | Jump to Content | cloud.google.com | Retrieved by this lane's web search. | |
| p22 | SPACE Metrics Framework for Developers Explained (2025 Edition) | LinearB Blog | linearb.io | April 18, 2025 | Retrieved by this lane's web search. |
| p23 | DevOps Metrics: DORA Metrics, SPACE Framework, DevEx - Travis CI | travis-ci.com | May 30, 2025 | Retrieved by this lane's web search. |
| p24 | Unlocking Engineering Excellence: DORA, SPACE, DevEX, and Beyond | by Soulaiman Ghanem | Code Factory Berlin | Medium | medium.com | February 1, 2025 | Retrieved by this lane's web search. |
| p25 | Harnessing the Potential of Gen-AI Coding Assistants in Public Sector Software Development | arxiv.org | Retrieved by this lane's web search. |
Academic & arXiv
Blogs & Independent Thinkers
| ID | Title | Outlet | Date | Significance |
|---|---|---|---|---|
| b1 | DORA 2025: Measuring Software Delivery After AI | RedMonk | 2025-12 | Follow-up independent analysis tracking the 2025 reversal in AI-throughput correlation and continued instability correlation, with direct quotes from the DORA report |
| b2 | On METR's AI Coding RCT | LessWrong | 2025-07 | LessWrong community analysis testing the METR RCT's generalisability, weighing the experience-with-tool hypothesis against the raw slowdown finding |
| b3 | We are Changing our Developer Productivity Experiment Design | METR (Substack) | 2026-02 | METR's own Substack admission that a 2026 follow-up study produced an unreliable signal due to participation bias, an important caveat on replication |
| b4 | AI coders think they're 20% faster - but they're actually 19% slower | Pivot to AI | 2025-07 | Skeptical independent blog treatment of the METR RCT that quotes the study's finding that over half of AI suggestions were unusable |
| b5 | AI Coding Tools Slowed Down Top Developers, METR Study Finds | DROIDS (Substack) | 2025-07 | Substack newsletter coverage summarising the METR RCT's design and caveats, distinguishing greenfield/junior-developer scenarios from the tested case |
| b6 | Vibe engineering | Simon Willison's Weblog | 2025-10 | Simon Willison's influential independent framing of disciplined AI-assisted development as an amplifier of existing engineering practice, echoing DORA's amplifier thesis from hands-on practice rather than survey data |
| b7 | How to Misuse & Abuse DORA Metrics | GetDX Newsletter (Substack) | 2022-07 | Abi Noda's newsletter analysis of practitioner Bryan Finster's paper on DORA metrics becoming vanity radiators rather than improvement levers, a durable independent critique predating the AI wave |
| b8 | Introducing the DX Core 4 | GetDX Newsletter (Substack) | 2024-12 | Founder account of why DORA, SPACE and DevEx were combined into a fourth framework, directly addressing DORA's scope limitation to the commit-to-release window |
| b9 | DORA's latest research on AI impact | GetDX Newsletter (Substack) | 2025-05 | Practitioner newsletter interview with DORA's Nathen Harvey and Derek DeBellis unpacking the 2025 AI-impact findings shortly after release |
| b10 | METR's study on how AI affects developer productivity | GetDX Newsletter (Substack) | 2025-07 | Independent newsletter summary of the METR RCT methodology and the five factors METR identified as contributing to the slowdown |
| b11 | DORA, SPACE, DevEx, DX Core 4 | Jens Oliver Meiert (personal blog) | 2025-02 | Independent personal-blog critique noting that DevEx and DX Core 4 still lack the adoption and longitudinal data that give DORA and older frameworks their credibility |
| b12 | DORA Metrics: What the Research Actually Says | zbmowrey.com (personal blog) | 2025 | Independent practitioner blog compiling year-by-year DORA numbers (elite multipliers, AI adoption correlations, documentation multiplier) with explicit sourcing discipline |
| b13 | Your DORA Metrics Are Lying to You: The AI ROI Crisis Tearing Through Engineering Teams | Full Stack AI Engineer (Substack) | 2026-03 | Substack newsletter argument that AI adoption breaks the interpretability of deployment frequency, lead time and change failure rate individually, requiring supplementary signals |
| b14 | DORA Metrics in the AI Era: Still Enough? | Unblocked (company blog) | 2026-06 | Blog analysis contrasting the 2024 and 2025 DORA AI-adoption correlations and arguing the four keys need supplementing with a rework/churn signal and a trust/perception check |
| b15 | DORA AI Capabilities Model: AI makes good teams better | doubleSlash Blog | 2026-03 | Independent engineering consultancy blog dissecting the seven AI capabilities and cross-referencing them against a companion piece on why sprints lose relevance under fast AI-assisted delivery |
| b16 | DORA the Explorer - how to unpack AI capabilities paradoxes | Diginomica | 2026-02 | Independent enterprise-tech analyst commentary situating the 2025 AI Capabilities Model within DORA's broader research arc since 2023 |
| b17 | From Vibe Engineering to Continuous AI | Continue Blog | 2025-10 | Independent developer-tools blog extending Simon Willison's vibe engineering framing into an organisational practice model, illustrating how the term propagated through practitioner discourse |
| b18 | Mastering DORA Metrics in DevOps: A Guide to Optimizing Software Delivery | neverblink.ai | May 26, 2026 | Retrieved by this lane's web search. |
| b19 | AI Productivity Paradox: Real Developer ROI in 2025 | digitalapplied.com | December 26, 2025 | Retrieved by this lane's web search. |
| b20 | Research Update: Algorithmic vs. Holistic Evaluation - METR | metr.org | August 13, 2025 | Retrieved by this lane's web search. |
| b21 | Is AI Actually Making You Dumber and Slower? | ainewstoday.substack.com | Retrieved by this lane's web search. | |
| b22 | Agentic Engineering Patterns - Simon Willison's Newsletter | simonw.substack.com | February 27, 2026 | Retrieved by this lane's web search. |
| b23 | High Leverage | Ep. #9, The AI Coding Paradigm Shift with Simon Willison | Heavybit | heavybit.com | May 5, 2026 | Retrieved by this lane's web search. |
| b24 | Simon Willison on delivering AI generated code | Stefan Judis Web Development | stefanjudis.com | December 19, 2025 | Retrieved by this lane's web search. |
| b25 | Simon Willison’s Weblog | simonwillison.net | August 2, 2026 | Retrieved by this lane's web search. |
Financial Press
| ID | Title | Outlet | Date | Significance |
|---|---|---|---|---|
| f1 | AI slows down some experienced software developers, study finds | Reuters (via Investing.com) | 2025-07 | Reuters wire report on the METR randomised controlled trial, the most rigorous causal evidence yet that AI coding assistance can increase task completion time despite near-universal developer belief in speedups. |
| f2 | We are Changing our Developer Productivity Experiment Design | METR | 2026-02 | METR's own follow-up disclosing that its second-wave experiment produced an unreliable signal due to selection effects, a methodological caveat rarely reported alongside the original 19 percent headline. |
| f3 | Announcing the 2025 DORA Report | Google Cloud Blog | 2025-09 | Google Cloud's official summary of the 2025 report, based on nearly 5,000 respondents and 100+ hours of interviews, introducing the AI Capabilities Model and the amplifier framing. |
| f4 | How are developers using AI? Inside Google's 2025 DORA report | Google Blog | 2025-09 | Google's own framing that AI adoption alone is insufficient and that a blend of technical and cultural capabilities determines outcomes, central to the accelerator vs amplifier distinction. |
| f5 | DORA | DORA 2025: Year in review | DORA | 2025 | DORA's own recap distinguishing the September State of AI-assisted Software Development report from the December AI Capabilities Model companion, useful for tracking methodology and publication cadence. |
| f6 | Sources: Cursor in talks to raise $2B+ at $50B valuation as enterprise growth surges | TechCrunch (citing Bloomberg) | 2026-04 | TechCrunch report citing Bloomberg on Cursor/Anysphere's revenue trajectory (over 2 billion in annualised revenue by February 2026, forecasting over 6 billion by end-2026), illustrating the scale of capital betting on AI-coding productivity claims. |
| f7 | Cursor's Anysphere nabs $9.9B valuation, soars past $500M ARR | TechCrunch (citing Bloomberg) | 2025-06 | TechCrunch report citing Bloomberg's original reporting on Anysphere's 900 million dollar raise, documenting investor conviction in AI-coding tools well ahead of independent productivity verification. |
| f8 | Anysphere, maker of AI coding tool Cursor, confirms $900M raise led by Thrive Capital | Crunchbase News | 2025-06 | Crunchbase News piece corroborating Bloomberg's description of Anysphere as "the fastest growing startup ever", tying the AI-coding capital story to specific investor names. |
| f9 | AI in Software Development: Productivity at the Cost of Code Quality? | DevOps.com (citing Wall Street Journal) | 2025-12 | Trade coverage quoting the Wall Street Journal's interview with MIT's Armando Solar-Lezama comparing AI coding assistance to a credit card that lets teams "accumulate technical debt", linking financial-press framing to DevOps quality concerns. |
| f10 | How AI is already shaking up the job market | Morning Brew | 2025-05 | Reports TD Bank's comments to the Wall Street Journal that it now screens for abstract coding and prompt-engineering ability, and Salesforce's claim of a 30% productivity boost behind its engineering hiring freeze. |
| f11 | AI's productivity paradox | AOL (Business Insider) | 2026-06 | Cites Moody's chief economist Mark Zandi arguing AI's productivity lift will emerge gradually rather than immediately, framing the gap between capital markets' AI enthusiasm and measured output. |
| f12 | The Economics of Intelligence | Citadel Securities | 2026-05 | Citadel Securities analysis applying the Jevons Paradox to software engineering, arguing cheaper AI-assisted coding could expand rather than shrink developer demand, a market-structure argument financial analysts use to explain flat headcount despite productivity claims. |
| f13 | Research: Quantifying GitHub Copilot's impact in the enterprise with Accenture | GitHub Blog | 2024-05 | Primary vendor-commissioned study (GitHub/Accenture) on enterprise Copilot telemetry, useful as a vendor claim requiring independent corroboration against DORA and METR findings. |
| f14 | AI is eroding code quality states new in-depth report | DevClass | 2025-02 | devclass coverage of GitClear's findings that duplicated code blocks rose eightfold in 2024, providing an independent technology-press read on the code-quality telemetry debate. |
| f15 | RDEL #112: What's AI's impact on software delivery performance? (2025 DORA Report) | Research-Driven Engineering Leadership (Substack) | 2025 | Practitioner analysis quantifying DORA's Bayesian modelling approach (nearly 5,000 professionals, 100+ hours of interviews) and its finding that AI adoption improved throughput but continued to increase instability in 2025. |
| f16 | How Generative AI Is Changing Software Development: Key Insights from the DORA Report | opslevel.com | Retrieved by this lane's web search. | |
| f17 | Top takeaways from the 2024 DORA Report – sponsored by Gearset! | DevOps Launchpad | devopslaunchpad.com | October 23, 2024 | Retrieved by this lane's web search. |
| f18 | DORA 2024: AI Is Speeding Up Delivery and Hurting Stability at the Same Time | Ali Safari | alisafari.space | May 14, 2026 | Retrieved by this lane's web search. |
| f19 | METR on X: "We ran a randomized controlled trial to see how much AI coding tools speed up experienced open-source developers. The results surprised us: Developers thought they were 20% faster with AI tools, but they were actually 19% slower when they had access to AI than when they didn't." / X | x.com | July 10, 2025 | Retrieved by this lane's web search. |
| f20 | My Participation in the METR AI Productivity Study | Domenic Denicola | domenic.me | July 15, 2025 | Retrieved by this lane's web search. |
| f21 | IACDM: Interactive Adversarial Convergence Development Methodology -- A Structured Framework for AI-Assisted Software Development | arxiv.org | Retrieved by this lane's web search. | |
| f22 | A METR Study Reveals that AI Slows Down Experienced Developers - ActuIA | actuia.com | July 16, 2025 | Retrieved by this lane's web search. |
| f23 | AI Coding Statistics - Adoption, Productivity & Market Metrics | getpanto.ai | August 7, 2026 | Retrieved by this lane's web search. |
Frontier Lab & Model News
| ID | Title | Outlet | Date | Significance |
|---|---|---|---|---|
| t1 | Details about METR's evaluation of OpenAI GPT-5 | METR | 2025-08 | Primary METR pre-deployment evaluation report giving GPT-5's autonomous coding time-horizon figures used to benchmark frontier capability growth against DORA-era instability findings. |
| t2 | METR's GPT-4.5 pre-deployment evaluations | METR | 2025-02 | Earlier METR evaluation establishing the time-horizon methodology later applied to GPT-5 and subsequent models, showing the capability trendline behind AI-coding adoption. |
| t3 | METR's Evaluation of GPT-5 | Alignment Forum | 2025-08 | Cross-posted METR researcher commentary giving additional statistical detail (bootstrap comparisons, confidence intervals) on the GPT-5 autonomy evaluation. |
| t4 | OpenAI GPT-5 System Card | OpenAI (arXiv mirror) | 2025 | Official system card documenting METR's external evaluation conclusions on GPT-5-thinking's autonomy and AI R&D speed-up risk, the primary-source counterpart to METR's own blog report. |
| t5 | GPT-5.6 Preview System Card - External Evaluations for AI Self-Improvement, METR | OpenAI Deployment Safety Hub | 2026 | OpenAI's Deployment Safety Hub documentation of continued METR involvement in evaluating successive GPT-5.x releases for autonomy and destructive-action risks. |
| t6 | System Card: Claude Opus 5 | Anthropic | 2026-07 | Anthropic's system card giving its own AI R&D capability index (ECI) comparison for Claude Opus 5, the frontier-lab counterpart to METR's independent time-horizon tracking. |
| t7 | 2025 DORA State of AI Assisted Software Development Report | intelligence.theregister.com | Retrieved by this lane's web search. | |
| t8 | AI Dev: The 2024 DORA Report Reviewed | by Julian B | Medium | medium.com | November 27, 2024 | Retrieved by this lane's web search. |
| t9 | DevOps in 2024: Essential Insights from the DORA Report | Upsun | upsun.com | November 19, 2025 | Retrieved by this lane's web search. |
| t10 | AI Writes 41% of Code: The Churn and Tech-Debt Data | particula.tech | June 5, 2026 | Retrieved by this lane's web search. |
| t11 | GPT-5.3 and Claude Opus 4.6: More System Card Shenanigans | ignorance.ai | February 11, 2026 | Retrieved by this lane's web search. |
| t12 | Understanding The 4 DORA Metrics And Top Findings From 2024/25 DORA Report | Octopus Deploy | octopus.com | Retrieved by this lane's web search. | |
| t13 | What are the DORA metrics? | dynatrace.com | October 17, 2025 | Retrieved by this lane's web search. |
| t14 | Understanding DORA Metrics for DevOps Performance | oneuptime.com | February 20, 2026 | Retrieved by this lane's web search. |
VC & Analyst Reports
| ID | Title | Outlet | Date | Significance |
|---|---|---|---|---|
| v1 | Gartner Magic Quadrant for AI Code Assistants | Gartner | 2025-09 | Gartner's primary 2025 vendor evaluation of 14 AI code-assistant providers, the analyst-radar reference point for this market. |
| v2 | The AI software development market map | CB Insights | 2025-10 | CB Insights market map of code-generation, review and testing vendors, with ARR growth figures for Anysphere/Cursor, Replit and Lovable. |
| v3 | The AI agent market map | CB Insights | 2026-03 | CB Insights' broader agent market map identifying software development as one of the two most-traction AI agent categories. |
| v4 | Developer Tools 2.0 | Sequoia Capital | 2023-05 | Sequoia's foundational 2023 investment thesis on Copilot-style AI coding tools, illustrating the early accelerator framing later complicated by DORA/METR data. |
| v5 | Services: The New Software | Sequoia Capital | 2026-03 | Sequoia's 2026 thesis pivot toward AI 'autopilots' replacing labour budgets rather than just tooling, showing the firm's thesis evolution beyond simple coding-assistant acceleration. |
| v6 | Unleash developer productivity with generative AI | McKinsey & Company | 2023-06 | McKinsey's original 2023 lab study of 40+ developers finding productivity and flow gains, a precursor to later organisation-level scepticism. |
| v7 | AI in software development: boosting productivity | McKinsey & Company | 2025-12 | McKinsey's 2025-2026 interview citing Jellyfish vendor data on adoption thresholds and productivity gains, illustrating vendor-sourced numbers entering consultancy commentary. |
| v8 | From Pilots to Payoff: Generative AI in Software Development | Bain & Company | 2025-09 | Bain's primary 2025 Technology Report finding only 10-15% productivity gains from AI coding assistants and arguing for lifecycle-wide redesign. |
| v9 | Technology Report 2025 - Technology Industry Trends | Bain & Company | 2025-09 | Bain's topic hub for its sixth annual Technology Report, framing AI productivity gains against required process change. |
| v10 | AI coding hype overblown, Bain shrugs | The Register | 2025-09 | Independent press summary of Bain's 2025 findings on low developer adoption and modest 10-15% productivity gains. |
| v11 | Predictions 2025: GenAI Reality Bites Back For Software Developers | Forrester | 2024-10 | Forrester's 2025 predictions, including that developers spend only ~24% of time coding and a prediction of a failed 50%-developer-replacement attempt. |
| v12 | What AI-Enhanced Software Development Means For Technology Executives | Forrester | 2025-09 | Forrester's 2025 developer-survey commentary on uneven AI-assistant adoption across SDLC stages ('TuringBots'). |
| v13 | AI Is Amplifying Software Engineering Performance, Says the 2025 DORA Report | InfoQ | 2026-03 | Independent trade-press synthesis of the 2025 DORA report's throughput-up/stability-down finding and the amplifier framing. |
| v14 | Enterprise Tech Investments & Team Overview | Andreessen Horowitz (a16z) | 2025 | a16z's enterprise AI investment hub, tracking enterprise gen-AI spend growth and top AI-native productivity companies relevant to market sizing. |
| v15 | Productivity and Pitfalls in AI Coding DORA 2025 | by Tamanna | Medium | medium.com | September 29, 2025 | Retrieved by this lane's web search. |
| v16 | AI Won’t Fix Broken Systems: Lessons from the 2025 DORA Report - Aviator Blog | aviator.co | June 7, 2026 | Retrieved by this lane's web search. |
| v17 | DORA Metrics Engineering Effectiveness: AI Impact in 2026 | blog.exceeds.ai | July 30, 2026 | Retrieved by this lane's web search. |
| v18 | DORA Report 2025: How AI Adoption Shapes DevOps and Software Teams - Opsera | opsera.ai | November 5, 2025 | Retrieved by this lane's web search. |
| v19 | 2025 DORA State of AI-assisted Software Development Report | cloud.google.com | Retrieved by this lane's web search. | |
| v20 | McKinsey is Still Talking about Eng Prod - Here's Why | faros.ai | February 26, 2026 | Retrieved by this lane's web search. |
| v21 | The Productivity Paradox: Why AI Developers Feel Faster But Deliver Slower | by Eran Swears | Medium | medium.com | December 9, 2025 | Retrieved by this lane's web search. |
| v22 | McKinsey State of AI 2025: What It Means for Engineering Leaders | colabsoftware.com | April 10, 2026 | Retrieved by this lane's web search. |
| v23 | McKinsey & Company: Measuring AI Developer Productivity & Attrition Risk 2026 | codeninety.com | February 18, 2026 | Retrieved by this lane's web search. |