Research · Financial Press

Back to sweep

Research sweep · deep · 2016 – 2026

Software Project Outcomes Over Time, and Is AI Improving Them

Software and IT project outcomes over time from July 2016 to July 2026 (biased to the latest year): baseline success/challenged/failure ranges across trusted longitudinal datasets (Standish CHAOS, Oxford Global Projects and Flyvbjerg, McKinsey, PMI Pulse), the drivers of outcomes (complexity, budget, domain, delivery methodology, sourcing), build-vs-buy and low-code/no-code commissioning choices (OutSystems, Mendix, Power Platform, Retool), and whether AI-assisted development (Copilot, Cursor, Claude Code, DORA and METR evidence) is yet improving delivery outcomes on traditional software projects that build ordinary systems rather than AI products.

  • Claude Fable 5
  • financial
  • frontier
  • academic
  • vc
  • blogs
  • tech

Synthesised 2026-07-24

Narrative

Financial-press coverage of software delivery outcomes clusters around three storylines: the durability of grim baseline statistics, high-profile public-sector cancellations, and a sharpening Wall Street debate over whether AI coding tools are actually paying off. The foundational quantitative source remains McKinsey's Oxford collaboration on 5,400 large IT projects, which found that on average large IT projects run 45 percent over budget and 7 percent over time, while delivering 56 percent less value than predicted, with software projects carrying the highest risk of cost and schedule overruns. That single 2012 dataset still anchors nearly every subsequent citation of "large IT project" failure rates across consultancies and trade press, which is itself a methodological concern: independent academic scrutiny of the Standish CHAOS figures (Eveleens and Verhoef, cited in the peer-reviewed literature) found that Standish's successful and challenged project results are indeed meaningless for benchmarking, even though the directional findings, that large projects fail far more than small ones, have been corroborated by separate McKinsey, BCG and PMI surveys.

Reuters' August 2025 investigation into the Pentagon's cancellation of two nearly finished Navy and Air Force HR software systems, built by Accenture and Oracle over 12 years for more than 800 million dollars, is the clearest financial-press case study of large-scale public-sector software failure in the current period. Reuters reported that both systems had cleared multiple independent technical reviews and an April status update had described the Air Force project as "on track," yet Pentagon leadership issued a "strategic pause" before scrapping the work in favour of new bids to Salesforce and Palantir, prompting Democratic lawmakers to call the reversal a "costly do-over." This episode illustrates a recurring theme in the public-sector literature: cancellations driven by governance and procurement politics rather than technical failure, with sunk costs approaching completion.

On the AI-assisted development question, Bloomberg Businessweek's February 2026 feature "AI Coding Agents Like Claude Code Are Fueling a Productivity Panic in Tech" captures the shift in mood a year after "vibe coding" entered the vocabulary: the promise of easier development instead touched off a high-pressure race to build software at any cost. This sits alongside Bloomberg's earnings coverage showing Wall Street turning skeptical of the hundreds of billions of dollars Big Tech is spending on AI, with investors demanding evidence of monetisation from Microsoft, Meta, Alphabet and Amazon before rewarding further capital expenditure. The rigorous instrumented evidence remains thin and contested: METR's randomized controlled trial of 16 experienced open-source developers found that allowing AI actually increased completion time by 19 percent even though developers had forecast a 24 percent speed-up beforehand, a result METR itself later cautioned may be an unreliable signal given participant self-selection bias in follow-up experiments. Google's 2025 DORA report, surveying nearly 5,000 technology professionals, found AI adoption surged to 90 percent but concluded AI does not fix a team, it amplifies what is already there, meaning strong teams improve further while struggling teams see existing problems intensified.

Enterprise commissioning coverage shows the low-code and AI-coding markets both scaling fast but without independent evidence resolving the build-versus-buy debate. Gartner's Magic Quadrant methodology and market-sizing estimates (cited across trade press rather than primary financial outlets) put the AI code-assistant market at roughly 3 to 3.5 billion dollars in 2025, while low-code platform forecasts from Gartner and IDC project the category will account for the majority of new enterprise applications by 2026. Reporting on enterprise AI cost governance, including CIO.com's coverage of AI cost overruns and Indeed Hiring Lab data (reported via Yahoo Finance) showing software developer job postings ticking up since Claude Code's 2025 release, point to a market still in the price-discovery phase rather than one with settled, audited productivity outcomes.


Sources

ID Title Outlet Date Significance
f1 AI Coding Agents Like Claude Code Are Fueling a Productivity Panic in Tech Bloomberg Businessweek 2026-02 Bloomberg Businessweek feature capturing the shift from AI-coding optimism to Wall Street skepticism about delivery pressure and ROI in the AI-assisted development era.
f2 Big Tech Earnings Land With 2026's AI Winners Still In Question Bloomberg 2026-01 Documents investor skepticism toward the hundreds of billions of dollars Big Tech is spending on AI with returns yet to materialize, framing the financial stakes behind AI-assisted development claims.
f3 Big Tech Earnings Show Split Between AI Trade Winners and Losers Bloomberg 2026-05 Shows differentiated market reaction to AI capital expenditure across Alphabet, Amazon, and Meta, relevant to enterprise investment flows into AI-assisted software delivery.
f4 Democrats decry move by Pentagon to pause $800 million in nearly done software projects Reuters (via syndication) 2025-08 Reuters original investigative reporting on the cancellation of two nearly complete Pentagon HR software systems after 12 years and $800m, a landmark recent public-sector IT failure case.
f5 Pentagon cancels Accenture, Oracle HR software after $800m spend People Matters (Reuters-sourced) 2025-08 Summarizes Reuters' reporting with direct quotes from Pentagon officials and contractors on the rationale for scrapping near-complete government software systems.
f6 Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity METR 2025-07 METR's randomized controlled trial found AI tooling increased task completion time by 19 percent despite developers forecasting a 24 percent speed-up, the most rigorous counter-evidence to AI productivity claims.
f7 We are Changing our Developer Productivity Experiment Design METR 2026-02 METR's own follow-up acknowledging that participant self-selection bias may make newer data an unreliable signal of AI's current productivity effect, a key caveat for the evidentiary base.
f8 Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (arXiv) arXiv 2025-07 Peer-reviewed preprint underlying the METR study, detailing the RCT design across 246 tasks in mature open-source repositories.
f9 Announcing the 2025 DORA Report Google Cloud 2025-09 Google Cloud's DORA report, based on nearly 5,000 technology professionals, articulates the amplifier thesis that AI magnifies existing team strengths and weaknesses rather than fixing organizational dysfunction.
f10 AI coding is now everywhere. But not everyone is convinced. MIT Technology Review 2025-12 MIT Technology Review synthesis distinguishing vendor-sponsored productivity studies (GitHub, Google, Microsoft showing 20-55% gains) from more mixed independent evidence.
f11 Delivering large-scale IT projects on time, on budget, and on value McKinsey & Company 2012-10 The foundational McKinsey-Oxford study of 5,400 IT projects underpinning nearly all subsequent large-IT-project failure statistics cited across consultancies and press.
f12 The Rise and Fall of the Chaos Report Figures ResearchGate (academic) 2025 Academic paper documenting the Eveleens and Verhoef critique that Standish's CHAOS success and challenge figures are methodologically unreliable for benchmarking, despite being the most widely cited industry statistic.
f13 Bent Flyvbjerg on Megaprojects: Why Big Things Go Wrong & How to Fix Them Thought Economics 2026-03 Interview with the Oxford professor behind the leading megaproject database (16,000+ projects), articulating the 'iron law' of over-budget, over-time, under-benefit outcomes and distinguishing reversible software projects from irreversible megaprojects.
f14 What you Should Know about Megaprojects and Why: An Overview Project Management Journal 2014 Flyvbjerg's peer-reviewed overview establishing the scale of global megaproject spending (6-9 trillion dollars annually) and formalizing the 'iron law of megaproject management' referenced across financial and policy press.
f15 Good news, software developers: Job postings are ticking up post-Claude Code Yahoo Finance / Indeed Hiring Lab 2026 Indeed Hiring Lab data reported via Yahoo Finance showing a 15 percent rise in US software developer job postings since Claude Code's 2025 release, relevant to labour-market effects of AI-assisted coding.
f16 AI cost overruns are adding up - with major implications for CIOs CIO 2025-10 CIO.com reporting on enterprise AI budget misestimation, showing a majority of organizations misjudge AI project costs by more than 10 percent, relevant to TCO concerns in build-vs-buy decisions.
f17 Nearly 7 in 10 firms report AI cost overruns CFO Dive 2026 CFO Dive coverage of enterprise AI spending governance failures, indicating financial-side scrutiny of AI project ROI parallel to software delivery outcome concerns.
f18 CIO, July 2026: AI is gaining momentum, cloud costs are rising, and cyber risks are on the increase Brandsit (citing Wall Street Journal) 2026-07 References Wall Street Journal reporting on AI agent sprawl at companies including Lyft, DaVita, GitLab and FICO, illustrating enterprise governance gaps as agentic coding scales.
f19 AI budgets soar, ROI still elusive Computerworld 2026-03 Computerworld analysis of the pattern of promising AI pilots followed by cost overruns and organizational skepticism, directly relevant to whether AI investment is translating into measurable delivery outcomes.
f20 Government Modernization Is Not a Software Problem Cicero Institute 2026-05 Policy analysis citing the Reuters Pentagon investigation and OECD research on how high-accountability public-sector cultures distort innovation and IT project outcomes.
f21 Developer Productivity Benchmarks 2026 | AI-Native Engineering Data larridin.com March 20, 2026 Retrieved by this lane's web search.
f22 Top 100 Developer Productivity Statistics with AI Tools 2026 index.dev Retrieved by this lane's web search.
f23 The AI Productivity Paradox Research Report faros.ai July 23, 2025 Retrieved by this lane's web search.
f24 93% of Developers Use AI. Why Is Productivity Only 10%? shiftmag.dev March 31, 2026 Retrieved by this lane's web search.
f25 AI Coding Statistics - Adoption, Productivity & Market Metrics getpanto.ai June 7, 2026 Retrieved by this lane's web search.

We use analytics cookies to understand site usage and improve the service. We do not use marketing cookies.