Research · Frontier Lab & Model News
Back to sweepResearch sweep · deep · 2016 – 2026
Software Project Outcomes Over Time, and Is AI Improving Them
Software and IT project outcomes over time from July 2016 to July 2026 (biased to the latest year): baseline success/challenged/failure ranges across trusted longitudinal datasets (Standish CHAOS, Oxford Global Projects and Flyvbjerg, McKinsey, PMI Pulse), the drivers of outcomes (complexity, budget, domain, delivery methodology, sourcing), build-vs-buy and low-code/no-code commissioning choices (OutSystems, Mendix, Power Platform, Retool), and whether AI-assisted development (Copilot, Cursor, Claude Code, DORA and METR evidence) is yet improving delivery outcomes on traditional software projects that build ordinary systems rather than AI products.
- Claude Fable 5
- financial
- frontier
- academic
- vc
- blogs
- tech
Synthesised 2026-07-24
Narrative
The most consequential evidence in this lane is METR's own longitudinal experimentation on AI coding productivity, which forms the empirical spine of the "does AI assistance help ordinary software delivery" debate. In July 2025 METR (Becker, Rush, Barnes and Rein) published a randomised controlled trial in which 16 experienced open-source developers completed 246 real issues in mature repositories they knew well, each task randomly assigned to allow or disallow AI tools (primarily Cursor Pro with Claude 3.5/3.7 Sonnet). The headline finding inverted expectations: developers forecast a 24% speedup, felt afterwards that they had been sped up by 20%, but were measured to be 19% slower when using AI. METR explicitly frames this as a snapshot of early-2025 capability in one setting, not a general verdict on AI coding tools.
METR's follow-up work shows how fast the measurement problem itself has moved. A second study launched in August 2025 with 57 developers and over 800 tasks could not be trusted as a productivity signal: METR's February 2026 post "We are Changing our Developer Productivity Experiment Design" reports that widespread adoption of agentic tools such as Claude Code and Codex introduced selection and behavioural-change effects (developers refusing to work without AI, changed task-completion patterns) that broke the original RCT design, even as METR states informally that developers are likely more sped up in early 2026 than in early 2025. METR's May 2026 survey of 349 technical workers found a median self-reported 1.4 to 2x change in value of work from AI, while explicitly flagging reasons to be sceptical of that magnitude, including task-substitution effects that inflate perceived speedups. Separately, METR's time-horizon research (the "task-completion time horizon" metric) shows frontier models' capability on software-engineering benchmark tasks doubling roughly every four to seven months since 2019, with some analysts noting the 2025 rate accelerated to roughly 3.5 months, but this is a capability benchmark, not a measure of delivery outcomes on real projects.
Google's DORA programme (Google Cloud, in partnership with GitHub and DevOps Research and Assessment) provides the other major instrumented data source. The 2024 DORA report was the first large-scale study to link AI tool usage to software delivery metrics, and multiple secondary write-ups (GetDX, Mezmo) describe its central and repeated finding that greater AI usage correlated with worse software delivery throughput and stability even as individual-level perceived productivity rose. The 2025 "State of AI-assisted Software Development" report reframes this as a systems and organisational-capability problem rather than a tools problem, introducing team-profile segmentation and a "Value Stream Management" framing to explain why local AI gains often fail to translate into organisational delivery improvement.
Code-level evidence from GitClear complements the survey and RCT data with structural signals of quality erosion. GitClear's analysis of 211 million changed lines of code (2020-2024) found that copy-pasted code exceeded refactored ("moved") code for the first time in 2024, that duplicated code blocks rose roughly eightfold, and that code churn (code reverted within two weeks) roughly doubled from a pre-AI baseline. GitClear's 2026 follow-up work found heavy AI users out-produce non-users by four to ten times in raw output, but attributes most of that gap to pre-existing differences between adopters rather than AI's causal effect, with heavy users showing a more modest 25% velocity gain relative to their own past selves. Anthropic's Economic Index, tracking Claude and Claude Code usage, documents a parallel shift from "augmentation" toward "automation" in coding workflows (79% of Claude Code conversations classified as automation versus 49% for Claude.ai), which is relevant context for how frontier labs characterise coding-agent adoption, though it is usage-pattern data rather than delivery-outcome evidence.