Research · Financial Press
Back to sweepResearch sweep · deep · 2016 – 2026
Software Project Outcomes Over Time, and Is AI Improving Them
Software and IT project outcomes over time from July 2016 to July 2026 (biased to the latest year): baseline success/challenged/failure ranges across trusted longitudinal datasets (Standish CHAOS, Oxford Global Projects and Flyvbjerg, McKinsey, PMI Pulse), the drivers of outcomes (complexity, budget, domain, delivery methodology, sourcing), build-vs-buy and low-code/no-code commissioning choices (OutSystems, Mendix, Power Platform, Retool), and whether AI-assisted development (Copilot, Cursor, Claude Code, DORA and METR evidence) is yet improving delivery outcomes on traditional software projects that build ordinary systems rather than AI products.
- Claude Fable 5
- financial
- frontier
- academic
- vc
- blogs
- tech
Synthesised 2026-07-24
Narrative
Financial-press coverage of software delivery outcomes clusters around three storylines: the durability of grim baseline statistics, high-profile public-sector cancellations, and a sharpening Wall Street debate over whether AI coding tools are actually paying off. The foundational quantitative source remains McKinsey's Oxford collaboration on 5,400 large IT projects, which found that on average large IT projects run 45 percent over budget and 7 percent over time, while delivering 56 percent less value than predicted, with software projects carrying the highest risk of cost and schedule overruns. That single 2012 dataset still anchors nearly every subsequent citation of "large IT project" failure rates across consultancies and trade press, which is itself a methodological concern: independent academic scrutiny of the Standish CHAOS figures (Eveleens and Verhoef, cited in the peer-reviewed literature) found that Standish's successful and challenged project results are indeed meaningless for benchmarking, even though the directional findings, that large projects fail far more than small ones, have been corroborated by separate McKinsey, BCG and PMI surveys.
Reuters' August 2025 investigation into the Pentagon's cancellation of two nearly finished Navy and Air Force HR software systems, built by Accenture and Oracle over 12 years for more than 800 million dollars, is the clearest financial-press case study of large-scale public-sector software failure in the current period. Reuters reported that both systems had cleared multiple independent technical reviews and an April status update had described the Air Force project as "on track," yet Pentagon leadership issued a "strategic pause" before scrapping the work in favour of new bids to Salesforce and Palantir, prompting Democratic lawmakers to call the reversal a "costly do-over." This episode illustrates a recurring theme in the public-sector literature: cancellations driven by governance and procurement politics rather than technical failure, with sunk costs approaching completion.
On the AI-assisted development question, Bloomberg Businessweek's February 2026 feature "AI Coding Agents Like Claude Code Are Fueling a Productivity Panic in Tech" captures the shift in mood a year after "vibe coding" entered the vocabulary: the promise of easier development instead touched off a high-pressure race to build software at any cost. This sits alongside Bloomberg's earnings coverage showing Wall Street turning skeptical of the hundreds of billions of dollars Big Tech is spending on AI, with investors demanding evidence of monetisation from Microsoft, Meta, Alphabet and Amazon before rewarding further capital expenditure. The rigorous instrumented evidence remains thin and contested: METR's randomized controlled trial of 16 experienced open-source developers found that allowing AI actually increased completion time by 19 percent even though developers had forecast a 24 percent speed-up beforehand, a result METR itself later cautioned may be an unreliable signal given participant self-selection bias in follow-up experiments. Google's 2025 DORA report, surveying nearly 5,000 technology professionals, found AI adoption surged to 90 percent but concluded AI does not fix a team, it amplifies what is already there, meaning strong teams improve further while struggling teams see existing problems intensified.
Enterprise commissioning coverage shows the low-code and AI-coding markets both scaling fast but without independent evidence resolving the build-versus-buy debate. Gartner's Magic Quadrant methodology and market-sizing estimates (cited across trade press rather than primary financial outlets) put the AI code-assistant market at roughly 3 to 3.5 billion dollars in 2025, while low-code platform forecasts from Gartner and IDC project the category will account for the majority of new enterprise applications by 2026. Reporting on enterprise AI cost governance, including CIO.com's coverage of AI cost overruns and Indeed Hiring Lab data (reported via Yahoo Finance) showing software developer job postings ticking up since Claude Code's 2025 release, point to a market still in the price-discovery phase rather than one with settled, audited productivity outcomes.
Sources
| ID | Title | Outlet | Date | Significance |
|---|---|---|---|---|
| f1 | AI Coding Agents Like Claude Code Are Fueling a Productivity Panic in Tech | Bloomberg Businessweek | 2026-02 | Bloomberg Businessweek feature capturing the shift from AI-coding optimism to Wall Street skepticism about delivery pressure and ROI in the AI-assisted development era. |
| f2 | Big Tech Earnings Land With 2026's AI Winners Still In Question | Bloomberg | 2026-01 | Documents investor skepticism toward the hundreds of billions of dollars Big Tech is spending on AI with returns yet to materialize, framing the financial stakes behind AI-assisted development claims. |
| f3 | Big Tech Earnings Show Split Between AI Trade Winners and Losers | Bloomberg | 2026-05 | Shows differentiated market reaction to AI capital expenditure across Alphabet, Amazon, and Meta, relevant to enterprise investment flows into AI-assisted software delivery. |
| f4 | Democrats decry move by Pentagon to pause $800 million in nearly done software projects | Reuters (via syndication) | 2025-08 | Reuters original investigative reporting on the cancellation of two nearly complete Pentagon HR software systems after 12 years and $800m, a landmark recent public-sector IT failure case. |
| f5 | Pentagon cancels Accenture, Oracle HR software after $800m spend | People Matters (Reuters-sourced) | 2025-08 | Summarizes Reuters' reporting with direct quotes from Pentagon officials and contractors on the rationale for scrapping near-complete government software systems. |
| f6 | Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity | METR | 2025-07 | METR's randomized controlled trial found AI tooling increased task completion time by 19 percent despite developers forecasting a 24 percent speed-up, the most rigorous counter-evidence to AI productivity claims. |
| f7 | We are Changing our Developer Productivity Experiment Design | METR | 2026-02 | METR's own follow-up acknowledging that participant self-selection bias may make newer data an unreliable signal of AI's current productivity effect, a key caveat for the evidentiary base. |
| f8 | Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (arXiv) | arXiv | 2025-07 | Peer-reviewed preprint underlying the METR study, detailing the RCT design across 246 tasks in mature open-source repositories. |
| f9 | Announcing the 2025 DORA Report | Google Cloud | 2025-09 | Google Cloud's DORA report, based on nearly 5,000 technology professionals, articulates the amplifier thesis that AI magnifies existing team strengths and weaknesses rather than fixing organizational dysfunction. |
| f10 | AI coding is now everywhere. But not everyone is convinced. | MIT Technology Review | 2025-12 | MIT Technology Review synthesis distinguishing vendor-sponsored productivity studies (GitHub, Google, Microsoft showing 20-55% gains) from more mixed independent evidence. |
| f11 | Delivering large-scale IT projects on time, on budget, and on value | McKinsey & Company | 2012-10 | The foundational McKinsey-Oxford study of 5,400 IT projects underpinning nearly all subsequent large-IT-project failure statistics cited across consultancies and press. |
| f12 | The Rise and Fall of the Chaos Report Figures | ResearchGate (academic) | 2025 | Academic paper documenting the Eveleens and Verhoef critique that Standish's CHAOS success and challenge figures are methodologically unreliable for benchmarking, despite being the most widely cited industry statistic. |
| f13 | Bent Flyvbjerg on Megaprojects: Why Big Things Go Wrong & How to Fix Them | Thought Economics | 2026-03 | Interview with the Oxford professor behind the leading megaproject database (16,000+ projects), articulating the 'iron law' of over-budget, over-time, under-benefit outcomes and distinguishing reversible software projects from irreversible megaprojects. |
| f14 | What you Should Know about Megaprojects and Why: An Overview | Project Management Journal | 2014 | Flyvbjerg's peer-reviewed overview establishing the scale of global megaproject spending (6-9 trillion dollars annually) and formalizing the 'iron law of megaproject management' referenced across financial and policy press. |
| f15 | Good news, software developers: Job postings are ticking up post-Claude Code | Yahoo Finance / Indeed Hiring Lab | 2026 | Indeed Hiring Lab data reported via Yahoo Finance showing a 15 percent rise in US software developer job postings since Claude Code's 2025 release, relevant to labour-market effects of AI-assisted coding. |
| f16 | AI cost overruns are adding up - with major implications for CIOs | CIO | 2025-10 | CIO.com reporting on enterprise AI budget misestimation, showing a majority of organizations misjudge AI project costs by more than 10 percent, relevant to TCO concerns in build-vs-buy decisions. |
| f17 | Nearly 7 in 10 firms report AI cost overruns | CFO Dive | 2026 | CFO Dive coverage of enterprise AI spending governance failures, indicating financial-side scrutiny of AI project ROI parallel to software delivery outcome concerns. |
| f18 | CIO, July 2026: AI is gaining momentum, cloud costs are rising, and cyber risks are on the increase | Brandsit (citing Wall Street Journal) | 2026-07 | References Wall Street Journal reporting on AI agent sprawl at companies including Lyft, DaVita, GitLab and FICO, illustrating enterprise governance gaps as agentic coding scales. |
| f19 | AI budgets soar, ROI still elusive | Computerworld | 2026-03 | Computerworld analysis of the pattern of promising AI pilots followed by cost overruns and organizational skepticism, directly relevant to whether AI investment is translating into measurable delivery outcomes. |
| f20 | Government Modernization Is Not a Software Problem | Cicero Institute | 2026-05 | Policy analysis citing the Reuters Pentagon investigation and OECD research on how high-accountability public-sector cultures distort innovation and IT project outcomes. |
| f21 | Developer Productivity Benchmarks 2026 | AI-Native Engineering Data | larridin.com | March 20, 2026 | Retrieved by this lane's web search. |
| f22 | Top 100 Developer Productivity Statistics with AI Tools 2026 | index.dev | Retrieved by this lane's web search. | |
| f23 | The AI Productivity Paradox Research Report | faros.ai | July 23, 2025 | Retrieved by this lane's web search. |
| f24 | 93% of Developers Use AI. Why Is Productivity Only 10%? | shiftmag.dev | March 31, 2026 | Retrieved by this lane's web search. |
| f25 | AI Coding Statistics - Adoption, Productivity & Market Metrics | getpanto.ai | June 7, 2026 | Retrieved by this lane's web search. |