Research · Academic & arXiv

Back to sweep

Research sweep · deep · 2016 – 2026

Software Project Outcomes Over Time, and Is AI Improving Them

Software and IT project outcomes over time from July 2016 to July 2026 (biased to the latest year): baseline success/challenged/failure ranges across trusted longitudinal datasets (Standish CHAOS, Oxford Global Projects and Flyvbjerg, McKinsey, PMI Pulse), the drivers of outcomes (complexity, budget, domain, delivery methodology, sourcing), build-vs-buy and low-code/no-code commissioning choices (OutSystems, Mendix, Power Platform, Retool), and whether AI-assisted development (Copilot, Cursor, Claude Code, DORA and METR evidence) is yet improving delivery outcomes on traditional software projects that build ordinary systems rather than AI products.

  • Claude Fable 5
  • financial
  • frontier
  • academic
  • vc
  • blogs
  • tech

Synthesised 2026-07-24

Narrative

The academic record on IT project outcomes converges on one durable finding and diverges sharply on everything else. The durable finding is that cost and schedule overruns on software and IT projects are fat-tailed rather than normally distributed. Flyvbjerg, Budzier, Lee, Keil, Lunn and Bester's 2022 Journal of Management Information Systems paper, analysing 5,392 IT projects, established that cost overruns follow a power-law distribution with a large mass of modest overruns and a small number of extreme blowouts, and proposed an interdependency-among-components mechanism to explain it. Their 2026 follow-up in the Journal of Cost Analysis and Parametrics tested 23 project types and found IT stands out with a 73-percentage-point gap between mean and median cost overrun, a signature of extreme fat tails, confirming IT projects carry unusually concentrated tail risk compared with other project categories. Companion arXiv working papers by Flyvbjerg and Budzier quantify this as roughly one in six large IT projects becoming a "black swan" with average cost overruns near 200% and schedule overruns of about 70%, a figure that underlies the widely cited McKinsey-Oxford 2012 study of 5,400-plus IT projects showing average overruns of 45% on cost and 7% on time while delivering 56% less value than predicted.

The Standish CHAOS figures, the most quoted baseline in practitioner literature, are treated far more skeptically in the peer-reviewed record. Eveleens and Verhoef's IEEE Software paper "The Rise and Fall of the Chaos Report Figures" argues Standish's success/challenged definitions rest solely on estimation accuracy, are one-sided, and are "meaningless for benchmarking" once real-world forecasting bias is accounted for. A 1304.4525 arXiv paper on public-sector IT project risk notes that academic scrutiny "debunked" the alarmist Standish-style findings after independent portfolio surveys failed to replicate an IT project crisis of that magnitude, and flags that using Standish's tripartite definitions can actively harm organisations that steer by them. Public-sector failure research (Springer meta-syntheses, case studies from Pakistan and New Zealand hospital IT) consistently finds governance, political stakeholders and underfunded "starved projects" (Flyvbjerg and Budzier document public-sector projects suffering average budget cuts of 75%) as failure modes distinct from, but compounding, the private-sector cost-overrun pattern.

On methodology, systematic literature reviews (Noteboom et al. 2021 HICSS; a 2025 Information and Software Technology SLR covering 208 articles from 1999-2024; a November 2025 arXiv PRISMA review of agile-to-hybrid evolution) repeatedly note that empirical evidence for agile improving project success is thin relative to its adoption, with one 2020 review explicitly stating there are "very few large-scale, empirical studies to support the contention that Agile methods can improve the likelihood of project success" despite widespread practitioner belief. Low-code/no-code research is similarly immature: a 2024 Journal of Systems and Software systematic review of 40 primary studies (2017-2023) and a 2026 ScienceDirect review of LCNC viability both find fragmented, largely qualitative evidence on adoption benefits and governance risk rather than head-to-head bespoke-versus-LCNC outcome data.

The most consequential recent academic development is METR's controlled experimental work on AI-assisted coding. Its July 2025 RCT (Becker, Rush, Barnes and Rein, arXiv 2507.09089), involving 16 experienced open-source developers completing 246 real tasks in mature repositories, found that allowing AI tools (chiefly Cursor Pro with Claude 3.5/3.7 Sonnet) made developers 19% slower, even though developers had forecast a 24% speedup and, after the fact, still believed they had been roughly 20% faster. METR has since revised its experimental design (task-level versus developer-level randomisation) and, per a February 2026 blog update, reports its second round of the same methodology on 57 developers now finds an 18% speedup, using late-2025 tools, illustrating how fast the underlying capability and hence the effect size is moving. METR's benchmark suite (HCAST, RE-Bench and SWAA, combined as METR-HRS) underpins its separate "time horizon" capability metric, showing autonomous task-completion length has doubled roughly every seven months since 2019, a capability trend distinct from, but frequently conflated with, delivery-outcome improvement. Independent longitudinal code-quality evidence from GitClear, based on analysis of over 200 million changed lines of code, documents a parallel maintainability cost: copy-pasted code overtook refactored ("moved") code for the first time in 2024, and code churn roughly doubled from a pre-AI baseline. Google's DORA 2025 State of AI-assisted Software Development report, drawing on nearly 5,000 respondents, frames AI as an "amplifier" that magnifies both high-performing and dysfunctional organisational patterns rather than improving delivery outcomes uniformly, a finding an arXiv 2026 paper on "the productivity-reliability paradox" tries to reconcile against GitHub's own RCT of code readability gains and DORA's simultaneous quality-gain/stability-loss result.


Sources

ID Title Outlet Date Significance
a1 The Empirical Reality of IT Project Cost Overruns: Discovering A Power-Law Distribution SSRN / Journal of Management Information Systems 2022-08 Foundational peer-reviewed (JMIS) demonstration that IT project cost overruns are fat-tailed/power-law distributed across 5,392 projects, the key academic counter to normal-distribution assumptions in project risk management.
a2 The Uniqueness of IT Cost Risk: A Cross-Group Comparison of 23 Project Types Journal of Cost Analysis and Parametrics (SAGE) 2026 2026 follow-up study testing whether IT cost overrun is more fat-tailed than 22 other project types, using 11,011 projects, extending Flyvbjerg et al.'s power-law findings.
a3 Overspend? Late? Failure? What the Data Say About IT Project Risk in the Public Sector arXiv 2013 arXiv working paper by Budzier and Flyvbjerg quantifying black-swan IT project risk and public-sector 'starved project' budget cuts, and critiquing Standish-style crisis narratives.
a4 Why Your IT Project Might Be Riskier Than You Think arXiv 2013 Companion arXiv paper (based on the Harvard Business Review 2011 article) quantifying the one-in-six black-swan IT project rate with 200% average cost overrun.
a5 The Oxford Olympics Study 2016: Cost and Cost Overrun at the Games arXiv 2016 Establishes Flyvbjerg's reference-class forecasting methodology later applied to IT project cost overrun research.
a6 The Rise and Fall of the Chaos Report Figures IEEE Software 2009 Canonical peer-reviewed critique arguing Standish CHAOS success/challenged definitions are methodologically meaningless for benchmarking.
a7 The Rise and Fall of the Chaos Report Figures (ResearchGate) ResearchGate 2009 Full-text version of the Eveleens and Verhoef critique showing political bias in IT forecasts undermines Standish benchmarking claims.
a8 The Standish Report: Does It Really Describe a Software Crisis? ResearchGate 2006 Academic paper assessing whether CHAOS-style figures represent a genuine software crisis or an artefact of biased sampling and definitions.
a9 Software Process Models and Analysis on Failure of Software Development Projects arXiv 2013 arXiv paper compiling multi-year Standish CHAOS trend data (1994-2009) alongside other failure-rate surveys such as TCS 2007.
a10 Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity arXiv / METR 2025-07 The core METR randomized controlled trial finding AI tools made experienced developers 19% slower despite believing they were faster, the most rigorous causal evidence on AI-assisted coding productivity.
a11 Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (METR blog) METR 2025-07 METR's own plain-language write-up including the July 2025 RCT design and a February 2026 update reporting a reversal to positive productivity with later AI tools.
a12 We are Changing our Developer Productivity Experiment Design METR 2026-02 METR's February 2026 methodological update describing a second RCT round (57 developers) and design changes addressing task-level versus developer-level randomization limitations.
a13 Measuring AI Ability to Complete Long Tasks METR 2025-03 METR's foundational time-horizon paper introducing the metric of AI task-completion length, showing a roughly 7-month doubling time, underpinning HCAST and RE-Bench based capability tracking.
a14 How Does Time Horizon Vary Across Domains? METR 2025-07 METR extension analysis testing whether the 7-month doubling trend generalizes beyond software/research tasks to other domains.
a15 Time Horizon 1.1 METR 2026-01 METR's 2026 update expanding the HCAST-derived task suite from 170 to 228 tasks, showing the evolving methodology behind autonomous capability benchmarking.
a16 Clarifying limitations of time horizon METR 2026-01 METR's own methodological caveats about benchmark selection bias and realism tradeoffs in its time-horizon capability metric.
a17 IACDM: Interactive Adversarial Convergence Development Methodology arXiv 2026 arXiv paper citing the METR RCT's three-way divergence between forecast, self-report, and measured productivity as motivation for a new AI-assisted development framework.
a18 The Productivity-Reliability Paradox: Specification-Driven Governance for AI-Augmented Software Development arXiv 2026 2026 arXiv paper attempting to reconcile conflicting empirical findings from GitHub's RCT, GitClear's longitudinal churn analysis, and DORA's 2024 report on AI code quality versus delivery stability.
a19 Adoption of low-code and no-code development: A systematic literature review and future research agenda Journal of Systems and Software 2024-11 2024 Journal of Systems and Software SLR of 40 primary studies (2017-2023) on LCNC and citizen development adoption, showing the academic evidence base is fragmented and largely qualitative.
a20 What does current research say about the viability of low-code development? A systematic literature review Journal of Systems and Software 2026-04 2026 ScienceDirect SLR covering January 2021-January 2025 publications assessing LCNC viability, usability and production-readiness evidence.
a21 An Empirical Study on Low-Code Programming using Traditional vs Large Language Model Support arXiv 2024-02 arXiv empirical comparison of traditional low-code development against LLM-assisted low-code approaches.
a22 The Evolution of Agile and Hybrid Project Management Methodologies: A Systematic Literature Review arXiv 2025-11 November 2025 arXiv PRISMA-guided SLR tracing the shift from pure agile to hybrid frameworks and identifying implementation challenges and success factors over 8 years of literature.
a23 Impact of Agile Methodology Use on Project Success in Organizations - A Systematic Literature Review ResearchGate 2020-12 SLR (30 primary studies from 1507 papers, 2008-2019) explicitly noting the scarcity of large-scale empirical evidence for agile's causal effect on project success despite its widespread adoption.
a24 A systematic literature review of agile software development projects Journal of Systems and Software 2025-03 2025 SLR of 208 articles (1999-2024) documenting persistent lack of convergence in agile software development research findings.
a25 Reasons for the Failure of Information Technology Projects in the Public Sector Springer 2021 Systematic literature review and meta-synthesis building a theoretical base for public-sector IT project failure factors, distinguishing governance and stakeholder drivers from private-sector patterns.

We use analytics cookies to understand site usage and improve the service. We do not use marketing cookies.