Research · Frontier Lab & Model News

Back to sweep

Research sweep · deep · 2016 – 2026

Software Project Outcomes Over Time, and Is AI Improving Them

Software and IT project outcomes over time from July 2016 to July 2026 (biased to the latest year): baseline success/challenged/failure ranges across trusted longitudinal datasets (Standish CHAOS, Oxford Global Projects and Flyvbjerg, McKinsey, PMI Pulse), the drivers of outcomes (complexity, budget, domain, delivery methodology, sourcing), build-vs-buy and low-code/no-code commissioning choices (OutSystems, Mendix, Power Platform, Retool), and whether AI-assisted development (Copilot, Cursor, Claude Code, DORA and METR evidence) is yet improving delivery outcomes on traditional software projects that build ordinary systems rather than AI products.

  • Claude Fable 5
  • financial
  • frontier
  • academic
  • vc
  • blogs
  • tech

Synthesised 2026-07-24

Narrative

The most consequential evidence in this lane is METR's own longitudinal experimentation on AI coding productivity, which forms the empirical spine of the "does AI assistance help ordinary software delivery" debate. In July 2025 METR (Becker, Rush, Barnes and Rein) published a randomised controlled trial in which 16 experienced open-source developers completed 246 real issues in mature repositories they knew well, each task randomly assigned to allow or disallow AI tools (primarily Cursor Pro with Claude 3.5/3.7 Sonnet). The headline finding inverted expectations: developers forecast a 24% speedup, felt afterwards that they had been sped up by 20%, but were measured to be 19% slower when using AI. METR explicitly frames this as a snapshot of early-2025 capability in one setting, not a general verdict on AI coding tools.

METR's follow-up work shows how fast the measurement problem itself has moved. A second study launched in August 2025 with 57 developers and over 800 tasks could not be trusted as a productivity signal: METR's February 2026 post "We are Changing our Developer Productivity Experiment Design" reports that widespread adoption of agentic tools such as Claude Code and Codex introduced selection and behavioural-change effects (developers refusing to work without AI, changed task-completion patterns) that broke the original RCT design, even as METR states informally that developers are likely more sped up in early 2026 than in early 2025. METR's May 2026 survey of 349 technical workers found a median self-reported 1.4 to 2x change in value of work from AI, while explicitly flagging reasons to be sceptical of that magnitude, including task-substitution effects that inflate perceived speedups. Separately, METR's time-horizon research (the "task-completion time horizon" metric) shows frontier models' capability on software-engineering benchmark tasks doubling roughly every four to seven months since 2019, with some analysts noting the 2025 rate accelerated to roughly 3.5 months, but this is a capability benchmark, not a measure of delivery outcomes on real projects.

Google's DORA programme (Google Cloud, in partnership with GitHub and DevOps Research and Assessment) provides the other major instrumented data source. The 2024 DORA report was the first large-scale study to link AI tool usage to software delivery metrics, and multiple secondary write-ups (GetDX, Mezmo) describe its central and repeated finding that greater AI usage correlated with worse software delivery throughput and stability even as individual-level perceived productivity rose. The 2025 "State of AI-assisted Software Development" report reframes this as a systems and organisational-capability problem rather than a tools problem, introducing team-profile segmentation and a "Value Stream Management" framing to explain why local AI gains often fail to translate into organisational delivery improvement.

Code-level evidence from GitClear complements the survey and RCT data with structural signals of quality erosion. GitClear's analysis of 211 million changed lines of code (2020-2024) found that copy-pasted code exceeded refactored ("moved") code for the first time in 2024, that duplicated code blocks rose roughly eightfold, and that code churn (code reverted within two weeks) roughly doubled from a pre-AI baseline. GitClear's 2026 follow-up work found heavy AI users out-produce non-users by four to ten times in raw output, but attributes most of that gap to pre-existing differences between adopters rather than AI's causal effect, with heavy users showing a more modest 25% velocity gain relative to their own past selves. Anthropic's Economic Index, tracking Claude and Claude Code usage, documents a parallel shift from "augmentation" toward "automation" in coding workflows (79% of Claude Code conversations classified as automation versus 49% for Claude.ai), which is relevant context for how frontier labs characterise coding-agent adoption, though it is usage-pattern data rather than delivery-outcome evidence.


Sources

ID Title Outlet Date Significance
t1 Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity arXiv / METR 2025-07 Primary METR RCT paper establishing the widely-cited 19% AI-induced slowdown finding among experienced developers.
t2 placeholder placeholder placeholder placeholder
t3 A randomized trial by METR found that experienced developers completed real coding tasks 19% slower when allowed to use AI tools - yet afterwards, they estimated on average that AI had made them 20% faster. - ScienceBlog.com scienceblog.com 2 weeks ago Retrieved by this lane's web search.
t4 Effect of AI on programming productivity – CMU S&DS Data Repository cmustatistics.github.io November 5, 2025 Retrieved by this lane's web search.
t5 AI Coding Tools Made Developers 19% Slower: METR Study | Let's Data Science letsdatascience.com March 27, 2026 Retrieved by this lane's web search.
t6 Research - METR metr.org Retrieved by this lane's web search.
t7 Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity - METR metr.org July 10, 2025 Retrieved by this lane's web search.
t8 My Participation in the METR AI Productivity Study | Domenic Denicola domenic.me July 15, 2025 Retrieved by this lane's web search.
t9 IACDM: Interactive Adversarial Convergence Development Methodology -- A Structured Framework for AI-Assisted Software Development arxiv.org Retrieved by this lane's web search.
t10 Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity arxiv.org Retrieved by this lane's web search.
t11 We are Changing our Developer Productivity Experiment Design - METR metr.org February 24, 2026 Retrieved by this lane's web search.
t12 ROI of AI-assisted Software Development dora.dev Retrieved by this lane's web search.
t13 DORA | State of AI-assisted Software Development 2025 dora.dev Retrieved by this lane's web search.
t14 DORA | Accelerate State of DevOps Report 2024 dora.dev Retrieved by this lane's web search.
t15 State of AI-assisted Software Development 2025 Platinum sponsors services.google.com Retrieved by this lane's web search.
t16 Highlights from the 2024 DORA State of DevOps Report getdx.com Retrieved by this lane's web search.
t17 2025 DORA State of AI Assisted Software Development | Google Cloud cloud.google.com Retrieved by this lane's web search.
t18 What the 2024 DORA Report Reveals About AI and DevOps mezmo.com Retrieved by this lane's web search.
t19 2025 DORA State of AI-assisted Software Development Report cloud.google.com Retrieved by this lane's web search.
t20 METR’s developer productivity research: 2026 update | Rob Bowley blog.robbowley.net April 4, 2026 Retrieved by this lane's web search.
t21 Measuring the Self-Reported Impact of Early-2026 AI on Technical Worker Productivity - METR metr.org May 11, 2026 Retrieved by this lane's web search.
t22 The AI Productivity Flip: What METR's 2026 Data Shows | PorkiCoder Blog porkicoder.com March 3, 2026 Retrieved by this lane's web search.
t23 AI Coding Productivity Study Data: What METR, McKinsey, and GitHub Found in 2026 | Value Add VC valueaddvc.com 2 weeks ago Retrieved by this lane's web search.
t24 AI Coding Productivity Paradox: 93% Adoption, 10% Gains - philippdubach.com philippdubach.com May 4, 2026 Retrieved by this lane's web search.
t25 GitHub Copilot Impact Analysis: ROI Metrics & Productivity blog.exceeds.ai April 15, 2026 Retrieved by this lane's web search.

We use analytics cookies to understand site usage and improve the service. We do not use marketing cookies.