Research · VC & Analyst Reports

Back to sweep

Research sweep · deep · 2016 – 2026

Software Project Outcomes Over Time, and Is AI Improving Them

Software and IT project outcomes over time from July 2016 to July 2026 (biased to the latest year): baseline success/challenged/failure ranges across trusted longitudinal datasets (Standish CHAOS, Oxford Global Projects and Flyvbjerg, McKinsey, PMI Pulse), the drivers of outcomes (complexity, budget, domain, delivery methodology, sourcing), build-vs-buy and low-code/no-code commissioning choices (OutSystems, Mendix, Power Platform, Retool), and whether AI-assisted development (Copilot, Cursor, Claude Code, DORA and METR evidence) is yet improving delivery outcomes on traditional software projects that build ordinary systems rather than AI products.

  • Claude Fable 5
  • financial
  • frontier
  • academic
  • vc
  • blogs
  • tech

Synthesised 2026-07-24

Narrative

Longitudinal software-project outcome data converges on a stable, uncomfortable baseline that has barely moved across three decades despite methodology change. Standish's CHAOS series, run since 1994, has become the most-cited but most-contested source: its founding figure was 16.2% of projects delivered on-time, on-budget, on-scope, with 31.1% cancelled and large-company projects delivering only 42% of planned features on average. By the 2020 edition, Standish reported 31% "successful," 50% "challenged" and 19% "failed," with small projects succeeding roughly 90% of the time versus under 10% for large ones, and government IT projects performing markedly worse. Independent academics (Eveleens and Verhoef's "Rise and Fall of the Chaos Report Figures") argue these figures are largely artefacts of shifting definitions and non-random sampling, and are "meaningless for benchmarking" as a result, a critique that has stuck through repeated citation cycles.

McKinsey's 2012 study with Oxford's BT Centre for Major Programme Management, based on more than 5,400 IT projects, found large IT projects (over $15 million) run 45% over budget and 7% over time on average while delivering 56% less value than predicted, with 17% classified as "black swan" projects with budget overruns exceeding 200%. Bent Flyvbjerg's Oxford Global Projects database extends this into fat-tail territory: IT projects show a power-law rather than normal distribution of cost overruns, a mean overrun near 73%, and, in the worst 18% of cases exceeding 50% overrun, a conditional tail expectation around 447%, worse than nuclear storage construction. PMI's Pulse of the Profession has moved toward a "value delivered" definition of success rather than the iron triangle, reporting a 73.8% average project performance rate in 2024 and finding no statistically significant difference in performance between agile, hybrid, predictive or remote/in-person delivery models, directly undercutting vendor claims that methodology choice alone drives outcomes.

On low-code/no-code, Gartner's Magic Quadrant for Enterprise Low-Code Application Platforms places OutSystems, Mendix, Microsoft Power Apps, ServiceNow, Appian and Salesforce as Leaders through the 2025 and 2026 cycles, and Gartner projects that by 2026 roughly 75% of new applications will be built on low-code platforms, with the market reaching an estimated $44.5 billion. Independent commentary increasingly flags a governance gap: citizen developers lack testing and security expertise, and analysts warn of a "2027-2028 technical debt reckoning" as ungoverned low-code proliferates, an important corrective to purely vendor-driven adoption-curve framing.

On AI-assisted coding, the strongest independent evidence is genuinely mixed and time-sensitive. METR's July 2025 randomised controlled trial of 16 experienced open-source developers found AI tools made them 19% slower on complex, familiar codebases, even though developers believed AI had sped them up by 20%; METR's own February 2026 update reports the same cohort's estimated effect flipping to an 18% speedup, while cautioning that growing non-participation bias (developers refusing to work without AI) makes the newer estimate unreliable. Google's DORA reports show a similar reversal at industry scale: the 2024 report found AI adoption correlated with lower throughput and stability, while the 2025 "State of AI-assisted Software Development" report (nearly 5,000 respondents) found AI adoption now correlates positively with throughput, product performance and individual effectiveness, but continues to correlate with worse delivery stability, a pattern DORA calls the "amplifier" effect: AI magnifies existing organisational strengths and dysfunctions rather than fixing them. GitClear's code-churn research across roughly 211 million lines (2020-2024) found refactoring collapsed from about 25% to under 10% of changed lines while duplicated code blocks rose eightfold and churn (code rewritten within two weeks) roughly doubled, suggesting throughput gains are partly offset by rework and maintainability costs that vendor-commissioned Forrester Total Economic Impact studies (376% claimed ROI for GitHub Copilot at scale) do not capture.


Sources

ID Title Outlet Date Significance
v1 The Rise and Fall of the Chaos Report Figures IEEE Software / ResearchGate 2010 Foundational academic critique arguing Standish's success/challenged figures are not valid for benchmarking due to political bias in IT forecasts.
v2 CHAOS Report on IT Project Outcomes OpenCommons 2025 Aggregates Standish CHAOS figures across editions (1994-2020), showing the 31% success/50% challenged/19% failed 2020 baseline and government IT underperformance.
v3 The Chaos Report - what makes projects successful LinkedIn / Clear Star Group 2025-02 Practitioner summary of Standish's final 2020 CHAOS methodology (50,000 projects, 104 factors) and the finding that team quality dominates other success factors.
v4 Why Your IT Project May Be Riskier than You Think Harvard Business Review / ResearchGate 2011 Flyvbjerg and Budzier's early (HBR-linked) analysis of 1,471 IT projects establishing the fat-tail cost-overrun framing later expanded in Oxford Global Projects work.
v5 Bent Flyvbjerg - University of Oxford profile and IT cost-risk research Academia.edu / University of Oxford 2009 Documents Flyvbjerg's finding that IT project cost risk has a fatter tail than any of 22 other project types studied, the basis of the Oxford Global Projects thesis.
v6 Software projects - the sting in the tail Longevitas 2023 Applies Flyvbjerg & Gardner (2023) Appendix A data to show IT projects have a 73% mean cost overrun and a 447% conditional tail expectation for the worst 18% of projects.
v7 Overspend? Late? Failure? What the Data Say About IT Project Risk in the Public Sector arXiv 2013 Oxford academic paper by Budzier and Flyvbjerg specifically examining public-sector IT project risk data.
v8 Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (coverage) Independent practitioner analysis 2025-07 METR's randomised controlled trial finding AI tools made experienced developers 19% slower on complex tasks despite developers believing they were faster, the most rigorous independent AI-coding productivity study to date.
v9 We are Changing our Developer Productivity Experiment Design METR 2026-02 METR's own follow-up acknowledging its updated experiment shows unreliable signal due to selection bias, and that the original 20% slowdown estimate does not straightforwardly generalise.
v10 What METR's Study Missed About AI Productivity in the Wild Faros AI 2026-02 Industry telemetry (10,000+ developers) contrasting METR's lab findings with real-world organisational data showing task throughput up but delivery speed unchanged, illustrating the individual-versus-organisational productivity gap.
v11 Announcing the 2025 DORA Report Google Cloud Blog 2025-09 Google Cloud's DORA team confirms AI now correlates positively with software delivery throughput but continues to correlate negatively with delivery stability, reversing the 2024 finding.
v12 DORA | Balancing AI tensions: Moving from AI adoption to effective SDLC use DORA (Google Cloud) 2026-03 Articulates DORA's 'amplifier' thesis: AI magnifies the strengths of high-performing organisations and the dysfunctions of struggling ones rather than fixing underlying capability gaps.
v13 DORA Report 2025 Key Takeaways: AI Impact on Dev Metrics Faros AI 2025-09 Telemetry-based follow-up documenting an 'Acceleration Whiplash' pattern: rising throughput accompanied by sharply worsening PR review time, bug rates and incident rates in 2026 data.
v14 State of AI-assisted Software Development 2025 DORA / Google Cloud 2025 Primary DORA report PDF documenting the reversal from the 2024 finding that more AI use worsened delivery stability and throughput.
v15 Investing in GitButler Andreessen Horowitz 2026-04 a16z investment thesis describing developer role shifting from 'programmer' to 'system architect and agent orchestrator' as the basis for infrastructure bets in the coding-agent era.
v16 Even a16z VCs say no one really knows what an AI agent is TechCrunch 2025-05 Illustrates the definitional uncertainty within a16z's own infrastructure investing team about AI agents, useful for noting recency bias and hype-cycle dynamics in VC framing.
v17 Delivering large-scale IT projects on time, on budget, and on value McKinsey & Company 2012 McKinsey's canonical study with Oxford's BT Centre (5,400+ IT projects) establishing the 45% budget overrun, 7% schedule overrun, 56% value shortfall, and 17% black-swan-project baseline still cited across the industry.
v18 PMI Talent Triangle: The Key to Project Management Success PMI 2025-03 Summarises PMI Pulse of the Profession 2024 findings that project success rates are statistically similar across remote, hybrid and in-person delivery models.
v19 Pulse of the Profession 2024: The Future of Project Work Project Management Institute 2024 Primary PMI report finding no statistically significant difference in project performance between agile, hybrid and predictive methodologies, directly countering vendor-driven agile-superiority claims.
v20 PMI's 2025 Pulse of the Profession report Project Management Institute 2025 Introduces PMI's redefinition of project success as 'delivered value that was worth the effort and expense' rather than pure iron-triangle metrics.
v21 Gartner Magic Quadrant for Low-Code Platforms for Enterprises (2026) ToolJet 2026-01 Documents Gartner's current Leaders quadrant (OutSystems, Mendix, Power Apps, ServiceNow, Appian, Salesforce) for the low-code commissioning decision landscape.
v22 Top 8 Low-Code, No-Code Platforms Reshaping Software Delivery in 2026 GEM Corporation 2026 Cites Gartner's $44.5 billion 2026 low-code market forecast and Forrester's 87% enterprise developer adoption figure alongside vendor-specific lock-in details (e.g. Mendix migration requiring rebuilds).
v23 Low-Code Hits $44.5B: Gartner 2026 Forecast Explained byteiota 2026-01 Independent analysis warning of a coming 'technical debt reckoning' from ungoverned citizen-developer low-code adoption, a corrective to purely vendor-optimistic adoption-curve framing.
v24 AI Copilot Code Quality: 2025 Data Suggests 4x Growth in Code Clones GitClear 2025-02 GitClear's instrumented analysis of 211 million lines of code (2020-2024) showing refactoring collapse and duplicated-code growth as AI coding tools scaled, key independent code-quality evidence.
v25 Report Summary: GitClear AI Code Quality Research 2025 jonas.rs 2025-02 Detailed breakdown of GitClear churn statistics (5.5% to 7.9% two-week churn, refactoring collapse from 24.1% to 9.5%) used across the industry as independent evidence of AI-era code quality trends.

We use analytics cookies to understand site usage and improve the service. We do not use marketing cookies.