Research · VC & Analyst Reports
Back to sweepResearch sweep · deep · 2016 – 2026
Software Project Outcomes Over Time, and Is AI Improving Them
Software and IT project outcomes over time from July 2016 to July 2026 (biased to the latest year): baseline success/challenged/failure ranges across trusted longitudinal datasets (Standish CHAOS, Oxford Global Projects and Flyvbjerg, McKinsey, PMI Pulse), the drivers of outcomes (complexity, budget, domain, delivery methodology, sourcing), build-vs-buy and low-code/no-code commissioning choices (OutSystems, Mendix, Power Platform, Retool), and whether AI-assisted development (Copilot, Cursor, Claude Code, DORA and METR evidence) is yet improving delivery outcomes on traditional software projects that build ordinary systems rather than AI products.
- Claude Fable 5
- financial
- frontier
- academic
- vc
- blogs
- tech
Synthesised 2026-07-24
Narrative
Longitudinal software-project outcome data converges on a stable, uncomfortable baseline that has barely moved across three decades despite methodology change. Standish's CHAOS series, run since 1994, has become the most-cited but most-contested source: its founding figure was 16.2% of projects delivered on-time, on-budget, on-scope, with 31.1% cancelled and large-company projects delivering only 42% of planned features on average. By the 2020 edition, Standish reported 31% "successful," 50% "challenged" and 19% "failed," with small projects succeeding roughly 90% of the time versus under 10% for large ones, and government IT projects performing markedly worse. Independent academics (Eveleens and Verhoef's "Rise and Fall of the Chaos Report Figures") argue these figures are largely artefacts of shifting definitions and non-random sampling, and are "meaningless for benchmarking" as a result, a critique that has stuck through repeated citation cycles.
McKinsey's 2012 study with Oxford's BT Centre for Major Programme Management, based on more than 5,400 IT projects, found large IT projects (over $15 million) run 45% over budget and 7% over time on average while delivering 56% less value than predicted, with 17% classified as "black swan" projects with budget overruns exceeding 200%. Bent Flyvbjerg's Oxford Global Projects database extends this into fat-tail territory: IT projects show a power-law rather than normal distribution of cost overruns, a mean overrun near 73%, and, in the worst 18% of cases exceeding 50% overrun, a conditional tail expectation around 447%, worse than nuclear storage construction. PMI's Pulse of the Profession has moved toward a "value delivered" definition of success rather than the iron triangle, reporting a 73.8% average project performance rate in 2024 and finding no statistically significant difference in performance between agile, hybrid, predictive or remote/in-person delivery models, directly undercutting vendor claims that methodology choice alone drives outcomes.
On low-code/no-code, Gartner's Magic Quadrant for Enterprise Low-Code Application Platforms places OutSystems, Mendix, Microsoft Power Apps, ServiceNow, Appian and Salesforce as Leaders through the 2025 and 2026 cycles, and Gartner projects that by 2026 roughly 75% of new applications will be built on low-code platforms, with the market reaching an estimated $44.5 billion. Independent commentary increasingly flags a governance gap: citizen developers lack testing and security expertise, and analysts warn of a "2027-2028 technical debt reckoning" as ungoverned low-code proliferates, an important corrective to purely vendor-driven adoption-curve framing.
On AI-assisted coding, the strongest independent evidence is genuinely mixed and time-sensitive. METR's July 2025 randomised controlled trial of 16 experienced open-source developers found AI tools made them 19% slower on complex, familiar codebases, even though developers believed AI had sped them up by 20%; METR's own February 2026 update reports the same cohort's estimated effect flipping to an 18% speedup, while cautioning that growing non-participation bias (developers refusing to work without AI) makes the newer estimate unreliable. Google's DORA reports show a similar reversal at industry scale: the 2024 report found AI adoption correlated with lower throughput and stability, while the 2025 "State of AI-assisted Software Development" report (nearly 5,000 respondents) found AI adoption now correlates positively with throughput, product performance and individual effectiveness, but continues to correlate with worse delivery stability, a pattern DORA calls the "amplifier" effect: AI magnifies existing organisational strengths and dysfunctions rather than fixing them. GitClear's code-churn research across roughly 211 million lines (2020-2024) found refactoring collapsed from about 25% to under 10% of changed lines while duplicated code blocks rose eightfold and churn (code rewritten within two weeks) roughly doubled, suggesting throughput gains are partly offset by rework and maintainability costs that vendor-commissioned Forrester Total Economic Impact studies (376% claimed ROI for GitHub Copilot at scale) do not capture.
Sources
| ID | Title | Outlet | Date | Significance |
|---|---|---|---|---|
| v1 | The Rise and Fall of the Chaos Report Figures | IEEE Software / ResearchGate | 2010 | Foundational academic critique arguing Standish's success/challenged figures are not valid for benchmarking due to political bias in IT forecasts. |
| v2 | CHAOS Report on IT Project Outcomes | OpenCommons | 2025 | Aggregates Standish CHAOS figures across editions (1994-2020), showing the 31% success/50% challenged/19% failed 2020 baseline and government IT underperformance. |
| v3 | The Chaos Report - what makes projects successful | LinkedIn / Clear Star Group | 2025-02 | Practitioner summary of Standish's final 2020 CHAOS methodology (50,000 projects, 104 factors) and the finding that team quality dominates other success factors. |
| v4 | Why Your IT Project May Be Riskier than You Think | Harvard Business Review / ResearchGate | 2011 | Flyvbjerg and Budzier's early (HBR-linked) analysis of 1,471 IT projects establishing the fat-tail cost-overrun framing later expanded in Oxford Global Projects work. |
| v5 | Bent Flyvbjerg - University of Oxford profile and IT cost-risk research | Academia.edu / University of Oxford | 2009 | Documents Flyvbjerg's finding that IT project cost risk has a fatter tail than any of 22 other project types studied, the basis of the Oxford Global Projects thesis. |
| v6 | Software projects - the sting in the tail | Longevitas | 2023 | Applies Flyvbjerg & Gardner (2023) Appendix A data to show IT projects have a 73% mean cost overrun and a 447% conditional tail expectation for the worst 18% of projects. |
| v7 | Overspend? Late? Failure? What the Data Say About IT Project Risk in the Public Sector | arXiv | 2013 | Oxford academic paper by Budzier and Flyvbjerg specifically examining public-sector IT project risk data. |
| v8 | Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (coverage) | Independent practitioner analysis | 2025-07 | METR's randomised controlled trial finding AI tools made experienced developers 19% slower on complex tasks despite developers believing they were faster, the most rigorous independent AI-coding productivity study to date. |
| v9 | We are Changing our Developer Productivity Experiment Design | METR | 2026-02 | METR's own follow-up acknowledging its updated experiment shows unreliable signal due to selection bias, and that the original 20% slowdown estimate does not straightforwardly generalise. |
| v10 | What METR's Study Missed About AI Productivity in the Wild | Faros AI | 2026-02 | Industry telemetry (10,000+ developers) contrasting METR's lab findings with real-world organisational data showing task throughput up but delivery speed unchanged, illustrating the individual-versus-organisational productivity gap. |
| v11 | Announcing the 2025 DORA Report | Google Cloud Blog | 2025-09 | Google Cloud's DORA team confirms AI now correlates positively with software delivery throughput but continues to correlate negatively with delivery stability, reversing the 2024 finding. |
| v12 | DORA | Balancing AI tensions: Moving from AI adoption to effective SDLC use | DORA (Google Cloud) | 2026-03 | Articulates DORA's 'amplifier' thesis: AI magnifies the strengths of high-performing organisations and the dysfunctions of struggling ones rather than fixing underlying capability gaps. |
| v13 | DORA Report 2025 Key Takeaways: AI Impact on Dev Metrics | Faros AI | 2025-09 | Telemetry-based follow-up documenting an 'Acceleration Whiplash' pattern: rising throughput accompanied by sharply worsening PR review time, bug rates and incident rates in 2026 data. |
| v14 | State of AI-assisted Software Development 2025 | DORA / Google Cloud | 2025 | Primary DORA report PDF documenting the reversal from the 2024 finding that more AI use worsened delivery stability and throughput. |
| v15 | Investing in GitButler | Andreessen Horowitz | 2026-04 | a16z investment thesis describing developer role shifting from 'programmer' to 'system architect and agent orchestrator' as the basis for infrastructure bets in the coding-agent era. |
| v16 | Even a16z VCs say no one really knows what an AI agent is | TechCrunch | 2025-05 | Illustrates the definitional uncertainty within a16z's own infrastructure investing team about AI agents, useful for noting recency bias and hype-cycle dynamics in VC framing. |
| v17 | Delivering large-scale IT projects on time, on budget, and on value | McKinsey & Company | 2012 | McKinsey's canonical study with Oxford's BT Centre (5,400+ IT projects) establishing the 45% budget overrun, 7% schedule overrun, 56% value shortfall, and 17% black-swan-project baseline still cited across the industry. |
| v18 | PMI Talent Triangle: The Key to Project Management Success | PMI | 2025-03 | Summarises PMI Pulse of the Profession 2024 findings that project success rates are statistically similar across remote, hybrid and in-person delivery models. |
| v19 | Pulse of the Profession 2024: The Future of Project Work | Project Management Institute | 2024 | Primary PMI report finding no statistically significant difference in project performance between agile, hybrid and predictive methodologies, directly countering vendor-driven agile-superiority claims. |
| v20 | PMI's 2025 Pulse of the Profession report | Project Management Institute | 2025 | Introduces PMI's redefinition of project success as 'delivered value that was worth the effort and expense' rather than pure iron-triangle metrics. |
| v21 | Gartner Magic Quadrant for Low-Code Platforms for Enterprises (2026) | ToolJet | 2026-01 | Documents Gartner's current Leaders quadrant (OutSystems, Mendix, Power Apps, ServiceNow, Appian, Salesforce) for the low-code commissioning decision landscape. |
| v22 | Top 8 Low-Code, No-Code Platforms Reshaping Software Delivery in 2026 | GEM Corporation | 2026 | Cites Gartner's $44.5 billion 2026 low-code market forecast and Forrester's 87% enterprise developer adoption figure alongside vendor-specific lock-in details (e.g. Mendix migration requiring rebuilds). |
| v23 | Low-Code Hits $44.5B: Gartner 2026 Forecast Explained | byteiota | 2026-01 | Independent analysis warning of a coming 'technical debt reckoning' from ungoverned citizen-developer low-code adoption, a corrective to purely vendor-optimistic adoption-curve framing. |
| v24 | AI Copilot Code Quality: 2025 Data Suggests 4x Growth in Code Clones | GitClear | 2025-02 | GitClear's instrumented analysis of 211 million lines of code (2020-2024) showing refactoring collapse and duplicated-code growth as AI coding tools scaled, key independent code-quality evidence. |
| v25 | Report Summary: GitClear AI Code Quality Research 2025 | jonas.rs | 2025-02 | Detailed breakdown of GitClear churn statistics (5.5% to 7.9% two-week churn, refactoring collapse from 24.1% to 9.5%) used across the industry as independent evidence of AI-era code quality trends. |