Research · Software Project Outcomes Over Time, and Is AI Improving Them
Back to researchResearch sweep · deep · 2016 – 2026
Software Project Outcomes Over Time, and Is AI Improving Them
Software and IT project outcomes over time from July 2016 to July 2026 (biased to the latest year): baseline success/challenged/failure ranges across trusted longitudinal datasets (Standish CHAOS, Oxford Global Projects and Flyvbjerg, McKinsey, PMI Pulse), the drivers of outcomes (complexity, budget, domain, delivery methodology, sourcing), build-vs-buy and low-code/no-code commissioning choices (OutSystems, Mendix, Power Platform, Retool), and whether AI-assisted development (Copilot, Cursor, Claude Code, DORA and METR evidence) is yet improving delivery outcomes on traditional software projects that build ordinary systems rather than AI products.
Synthesised 2026-07-24
Software and IT Project Outcomes, 2016–2026: The Baselines Barely Moved, Then AI Arrived
Overview
Three decades of longitudinal measurement say roughly the same uncomfortable thing: most software projects of any size miss their targets, and the biggest ones miss catastrophically. Standish's CHAOS series began in 1994 with 16.2 percent of projects succeeding outright and 31.1 percent cancelled; by its 2020 edition the split was 31 percent successful, 50 percent challenged and 19 percent failed, with small projects succeeding roughly nine times as often as large ones. The McKinsey and Oxford study of more than 5,400 IT projects found large programmes running 45 percent over budget while delivering 56 percent less value than promised. Those numbers have anchored practitioner discourse ever since, despite sustained academic criticism of how they were produced. Sources: personal.utdallas.edu (n.d.) (↗); OpenCommons (2025) (↗); McKinsey & Company (2012) (↗)
The more rigorous story sits in Bent Flyvbjerg's Oxford work. His 2022 Journal of Management Information Systems paper on 5,392 IT projects showed cost overruns follow a power law, not a bell curve: a mass of modest misses plus a thin tail of extreme blowouts, with roughly one project in six becoming a "black swan" averaging near 200 percent cost overrun. A 2026 follow-up across 23 project types found IT uniquely fat-tailed, with a 73 percentage point gap between mean and median overrun. The practical implication is that average failure rates matter less than tail exposure, and tail exposure scales with project size. Sources: SSRN / Journal of Management Information Systems (2022) (↗); Journal of Cost Analysis and Parametrics (SAGE) (2026) (↗); arXiv (2013) (↗)
The defining shift of the past eighteen months is that the AI-assisted development question moved from vendor marketing to contested instrumented evidence, and the evidence refused to settle. METR's July 2025 randomised trial found experienced developers 19 percent slower with AI while believing themselves 20 percent faster. By February 2026 METR reported the effect had flipped to an 18 percent speedup with late-2025 agentic tools, while warning that ubiquitous adoption had broken its own experimental design. Google's DORA programme, meanwhile, moved from finding AI correlated with worse delivery throughput and stability in 2024 to a 2025 "amplifier" thesis: AI magnifies whatever organisational condition already exists. Sources: METR (2025) (↗); METR (2026) (↗); dora.dev (n.d.) (↗); Google Cloud (2025) (↗)
What has not yet happened is any connection between the two literatures. Nobody has published a longitudinal dataset linking AI tool adoption to CHAOS-style success and failure rates on ordinary business systems. The productivity debate runs on task-level trials, self-report surveys and code-quality telemetry; the outcomes literature runs on portfolio and programme data. The gap between them is the single most important evidentiary hole in this field.
Timeline
- Standish and McKinsey figures dominate practitioner baselines despite unresolved academic critique
- CHAOS 2020 edition reports 31/50/19 split with size the dominant driver
- GitClear's pre-AI code-quality baseline period begins
- GitHub Copilot launches, vendor-commissioned productivity studies set an optimistic narrative
- Flyvbjerg establishes power-law distribution of IT cost overruns
- ChatGPT triggers mass developer AI adoption
- Vendor case studies claim 40 to 55 percent task speedups on narrow work
- DORA links AI use to worse delivery throughput and stability
- GitClear finds copy-pasted code overtakes refactoring
- PMI shifts success definition toward value delivered
- METR RCT finds 19 percent slowdown against 24 percent forecast speedup
- Pentagon cancels 800 million dollars of near-finished HR systems
- DORA reports 90 percent adoption and the amplifier thesis
- Agentic tools (Claude Code, Cursor) become default workflow
- METR redesigns its experiment and reports the effect flipping positive
- Bloomberg names an AI productivity panic
- Flyvbjerg confirms IT as uniquely fat-tailed across 23 project types
- Gartner sizes low-code at 44.5 billion dollars
Key Findings
The headline baselines are less trustworthy than their citation count suggests. Eveleens and Verhoef's IEEE Software critique showed Standish's success definitions rest solely on estimation accuracy, are one-sided, and produce figures "meaningless for benchmarking"; Jørgensen and Moløkken made a similar case in 2006. Every lane in this sweep independently flagged the same problem: the most-quoted numbers come from a commercial product with unpublished sampling, while the more transparent Flyvbjerg data covers large programmes, not the median enterprise project. Sources: IEEE Software (2009) (↗); ResearchGate (2006) (↗); cs.vu.nl (n.d.) (↗)
Size and tail risk explain more variance than any other factor. Standish data puts small-project success near 90 percent against under 10 percent for the largest; Flyvbjerg's conditional tail expectation for the worst 18 percent of IT projects reaches roughly 447 percent overrun, worse than nuclear storage construction. The convergent lesson across academic and analyst lanes is that decomposition, not methodology choice, is the strongest known lever. Sources: OpenCommons (2025) (↗); Longevitas (2023) (↗); SSRN / Journal of Management Information Systems (2022) (↗)
Public-sector failure is real but over-documented, and often political rather than technical. The Pentagon's August 2025 cancellation of Navy and Air Force HR systems built by Accenture and Oracle over twelve years for more than 800 million dollars came after the systems passed independent technical reviews; work was rebid to Salesforce and Palantir. Academic work adds "starved projects" as a distinct public-sector failure mode, with average budget cuts of 75 percent, while audit trails (NAO, GAO) make government failures disproportionately visible relative to undisclosed private-sector ones. Sources: Reuters (via syndication) (2025) (↗); People Matters (Reuters-sourced) (2025) (↗); arXiv (2013) (↗); Springer (2021) (↗)
The evidence that methodology moves outcomes is remarkably weak. Systematic reviews from 2020 through 2025 repeatedly note "very few large-scale, empirical studies" supporting the claim that agile improves success likelihood, and PMI's 2024 Pulse found no statistically significant performance difference between agile, hybrid and predictive delivery. Effects that do appear are moderated by organisational maturity rather than method choice, which undercuts a large consulting and certification industry. Sources: ResearchGate (2020) (↗); arXiv (2025) (↗); Project Management Institute (2024) (↗)
METR's slowdown result remains the most rigorous single data point, and it has a date stamp. Sixteen experienced open-source developers, 246 real tasks in familiar repositories, random assignment: AI made them 19 percent slower while they forecast 24 percent faster and afterwards believed 20 percent faster. METR's February 2026 redesign note reports the second round (57 developers, 800-plus tasks) trending to an 18 percent speedup with late-2025 tools, but flags selection bias, including developers refusing to work without AI, as breaking the design. Any citation of either number without its date is misleading. Sources: arxiv.org (2025) (↗); metr.org (2025) (↗); metr.org (2026) (↗); blog.robbowley.net (2026) (↗)
DORA's amplifier thesis is the current best organising frame for AI and delivery. With adoption at 90 percent among nearly 5,000 respondents, the 2025 report found AI now correlates positively with throughput and individual effectiveness but still negatively with delivery stability. Strong teams compound gains; struggling teams see dysfunction intensify. The 2024-to-2025 reversal on throughput, from a Google-sponsored source, is either genuine capability improvement or partly reputational reframing; the lanes disagree. Sources: Google Cloud (2025) (↗); dora.dev (n.d.) (↗); getdx.com (n.d.) (↗); devops.com (2025) (↗)
Code-level telemetry shows a maintainability bill coming due. GitClear's analysis of 211 million changed lines (2020 to 2024) found copy-pasted code exceeding refactored code for the first time in 2024, duplicated blocks up roughly eightfold, and churn (code reverted within two weeks) roughly doubled from the pre-AI baseline. Its 2026 follow-up attributes most of the 4 to 10 times raw-output gap between heavy AI users and non-users to pre-existing differences, with a more modest 25 percent within-person velocity gain. Sources: GitClear (2025) (↗); jonas.rs (2025) (↗)
Low-code adoption claims run far ahead of independent outcome evidence. Gartner projects 75 percent of new applications on low-code platforms by 2026 and a 44.5 billion dollar market, but academic systematic reviews of LCNC (2024 and 2026) find fragmented, largely qualitative evidence with no head-to-head bespoke-versus-LCNC outcome data. Analysts warn of a governance gap and a possible technical debt reckoning around 2027 to 2028 as ungoverned citizen development accumulates. Sources: byteiota (2026) (↗); Journal of Systems and Software (2024) (↗); Journal of Systems and Software (2026) (↗); ToolJet (2026) (↗)
The perception gap is now itself a measured phenomenon. Developers in METR's trial misjudged their own speed by roughly 39 percentage points. METR's May 2026 survey of 349 technical workers found median self-reported value gains of 1.4 to 2 times, with METR itself flagging task-substitution effects that inflate perceived speedups. Bloomberg's February 2026 "productivity panic" framing captures the market consequence: capital markets demanding ROI evidence that self-report cannot supply. Sources: scienceblog.com (n.d.) (↗); metr.org (2026) (↗); Bloomberg Businessweek (2026) (↗)
Evidence & Data
The outcome baselines: Standish 1994 reported 16.2 percent success, 52.7 percent challenged, 31.1 percent cancelled; 2012 reported 37/42/21; 2020 reported 31/50/19. McKinsey-Oxford (5,400-plus projects) found 45 percent average cost overrun, 7 percent schedule overrun, 56 percent value shortfall, and 17 percent black swans exceeding 200 percent overrun. Flyvbjerg's database yields a mean IT overrun near 73 percent and a 447 percent conditional tail expectation in the worst 18 percent of cases. PMI's 2024 Pulse reports 73.8 percent average project performance under its value-based definition, an illustration of how much the definition moves the headline. Sources: personal.utdallas.edu (n.d.) (↗); OpenCommons (2025) (↗); McKinsey & Company (2012) (↗); Longevitas (2023) (↗); Project Management Institute (2024) (↗)
The AI evidence: METR's RCT measured a 19 percent slowdown (early 2025), then an unreliable 18 percent speedup signal (late 2025 tools). DORA 2025 recorded 90 percent AI adoption with persistent negative stability correlation. GitClear measured refactoring collapsing from about 25 percent to under 10 percent of changed lines. METR's separate time-horizon metric shows autonomous task length doubling every four to seven months since 2019, a capability trend frequently conflated with delivery improvement. Anthropic's Economic Index classifies 79 percent of Claude Code conversations as automation versus 49 percent for Claude.ai. On costs, nearly seven in ten firms report AI cost overruns, and Forrester's vendor-commissioned Copilot study claims 376 percent ROI, a figure best read as marketing. Sources: METR (2025) (↗); metr.org (2026) (↗); dora.dev (n.d.) (↗); GitClear (2025) (↗); METR (2025) (↗); CFO Dive (2026) (↗); blog.exceeds.ai (2026) (↗)
Signals & Tensions
Did the METR effect really flip? The 2026 positive estimate comes from a broken design, as METR itself admits. AI-optimists cite the flip as vindication; sceptics note the selection bias runs in the optimists' favour. Rob Bowley's independent read and Faros AI's field-data critique pull in opposite directions on the same numbers. Sources: metr.org (2026) (↗); blog.robbowley.net (2026) (↗); faros.ai (2026) (↗)
DORA's independence. The 2024 stability finding was credible precisely because it was unwelcome data from a Google-adjacent source. The 2025 shift to "systems problem, not tools problem" can be read as genuine insight or as reputational management, and the sweep's lanes split on which. Sources: mezmo.com (n.d.) (↗); dora.dev (n.d.) (↗)
Throughput versus stability is the emerging trade-off signature. DORA's persistent stability penalty, GitClear's churn doubling, and an arXiv "productivity-reliability paradox" paper all point the same way: raw output rises, downstream quality absorbs the cost. This is not yet consensus but it is the strongest cross-lane signal in the sweep. Sources: dora.dev (n.d.) (↗); GitClear (2025) (↗); arXiv (2026) (↗)
Capability curves versus delivery outcomes. METR's time-horizon doubling every four to seven months, possibly accelerating to 3.5 months in 2025, is a benchmark trend routinely conflated in press coverage with real project improvement, which the RCT evidence suggests lags well behind. Sources: METR (2025) (↗); METR (2026) (↗)
Public-sector visibility bias. Audit trails make government failures disproportionately documented, likely distorting popular perception of where failure risk concentrates. The Pentagon case adds a newer wrinkle: cancellation of technically passing projects for procurement-political reasons, sunk costs intact. Sources: Reuters (via syndication) (2025) (↗); Cicero Institute (2026) (↗)
Overhyped versus underreported. The 75 percent low-code adoption forecast is analyst rhetoric amplified by the platforms it describes. Underreported: the complete absence of independent LCNC-versus-bespoke outcome data, and rising developer job postings after Claude Code's release, which complicates the displacement narrative. Sources: ToolJet (2026) (↗); Yahoo Finance / Indeed Hiring Lab (2026) (↗)
Open Questions
-
Does AI adoption move CHAOS-style outcome rates at all? No published dataset yet links AI-assisted development to success, challenged or failure rates on traditional business-system delivery. Task-level speedups and portfolio-level outcomes remain unconnected literatures.
-
Can any clean causal measurement survive ubiquitous adoption? METR's own RCT design broke when developers refused to work without AI. If "AI versus no-AI" comparisons are no longer runnable, the field may be permanently limited to observational designs with heavy selection effects. Sources: metr.org (2026) (↗)
-
When does GitClear's maintainability debt convert into delivery failure? Duplication and churn are leading indicators; nobody has yet measured whether AI-era codebases fail, stall or cost more three years on. Sources: GitClear (2025) (↗)
-
Have baseline success rates actually improved since 2016, or have definitions moved? PMI's shift to value-based measurement and Standish's changing definitions make like-for-like comparison across the decade close to impossible, and no source in the sweep resolves it. Sources: IEEE Software (2009) (↗); Project Management Institute (2024) (↗)
-
Does AI help or harm by seniority and codebase age? METR's slowdown was measured on experts in mature, familiar repositories, arguably AI's worst case. Whether juniors on greenfield work see the mirror image is asserted everywhere and measured nowhere independently. Sources: getdx.com (2025) (↗)
-
Will the predicted low-code technical debt reckoning materialise? The 2027 to 2028 warning is plausible but currently forecast, not evidence, and post-mortems of LCNC total cost of ownership remain thin in the public record. Sources: byteiota (2026) (↗); Journal of Systems and Software (2026) (↗)
-
Does Flyvbjerg's fat-tail mechanism, interdependency among components, apply to AI-generated codebases? If AI increases coupling and duplication while accelerating output, tail risk on large programmes could worsen even as median velocity improves. Nobody has tested this. Sources: SSRN / Journal of Management Information Systems (2022) (↗)
The field's honest position in mid-2026: the old failure statistics were never as precise as their citation count implied, the new productivity statistics are moving too fast to pin down, and the one question commissioning organisations most need answered, whether AI-assisted teams ship ordinary systems more successfully, has not yet been measured by anyone.
![[sources-software-and-it-project-outcomes-over-time-from-ju]]
Sources
Summary: ↑ Back to summary
Financial Press
| ID | Title | Outlet | Date | Significance |
|---|---|---|---|---|
| f1 | AI Coding Agents Like Claude Code Are Fueling a Productivity Panic in Tech | Bloomberg Businessweek | 2026-02 | Bloomberg Businessweek feature capturing the shift from AI-coding optimism to Wall Street skepticism about delivery pressure and ROI in the AI-assisted development era. |
| f2 | Big Tech Earnings Land With 2026's AI Winners Still In Question | Bloomberg | 2026-01 | Documents investor skepticism toward the hundreds of billions of dollars Big Tech is spending on AI with returns yet to materialize, framing the financial stakes behind AI-assisted development claims. |
| f3 | Big Tech Earnings Show Split Between AI Trade Winners and Losers | Bloomberg | 2026-05 | Shows differentiated market reaction to AI capital expenditure across Alphabet, Amazon, and Meta, relevant to enterprise investment flows into AI-assisted software delivery. |
| f4 | Democrats decry move by Pentagon to pause $800 million in nearly done software projects | Reuters (via syndication) | 2025-08 | Reuters original investigative reporting on the cancellation of two nearly complete Pentagon HR software systems after 12 years and $800m, a landmark recent public-sector IT failure case. |
| f5 | Pentagon cancels Accenture, Oracle HR software after $800m spend | People Matters (Reuters-sourced) | 2025-08 | Summarizes Reuters' reporting with direct quotes from Pentagon officials and contractors on the rationale for scrapping near-complete government software systems. |
| f6 | Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity | METR | 2025-07 | METR's randomized controlled trial found AI tooling increased task completion time by 19 percent despite developers forecasting a 24 percent speed-up, the most rigorous counter-evidence to AI productivity claims. |
| f7 | We are Changing our Developer Productivity Experiment Design | METR | 2026-02 | METR's own follow-up acknowledging that participant self-selection bias may make newer data an unreliable signal of AI's current productivity effect, a key caveat for the evidentiary base. |
| f8 | Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (arXiv) | arXiv | 2025-07 | Peer-reviewed preprint underlying the METR study, detailing the RCT design across 246 tasks in mature open-source repositories. |
| f9 | Announcing the 2025 DORA Report | Google Cloud | 2025-09 | Google Cloud's DORA report, based on nearly 5,000 technology professionals, articulates the amplifier thesis that AI magnifies existing team strengths and weaknesses rather than fixing organizational dysfunction. |
| f10 | AI coding is now everywhere. But not everyone is convinced. | MIT Technology Review | 2025-12 | MIT Technology Review synthesis distinguishing vendor-sponsored productivity studies (GitHub, Google, Microsoft showing 20-55% gains) from more mixed independent evidence. |
| f11 | Delivering large-scale IT projects on time, on budget, and on value | McKinsey & Company | 2012-10 | The foundational McKinsey-Oxford study of 5,400 IT projects underpinning nearly all subsequent large-IT-project failure statistics cited across consultancies and press. |
| f12 | The Rise and Fall of the Chaos Report Figures | ResearchGate (academic) | 2025 | Academic paper documenting the Eveleens and Verhoef critique that Standish's CHAOS success and challenge figures are methodologically unreliable for benchmarking, despite being the most widely cited industry statistic. |
| f13 | Bent Flyvbjerg on Megaprojects: Why Big Things Go Wrong & How to Fix Them | Thought Economics | 2026-03 | Interview with the Oxford professor behind the leading megaproject database (16,000+ projects), articulating the 'iron law' of over-budget, over-time, under-benefit outcomes and distinguishing reversible software projects from irreversible megaprojects. |
| f14 | What you Should Know about Megaprojects and Why: An Overview | Project Management Journal | 2014 | Flyvbjerg's peer-reviewed overview establishing the scale of global megaproject spending (6-9 trillion dollars annually) and formalizing the 'iron law of megaproject management' referenced across financial and policy press. |
| f15 | Good news, software developers: Job postings are ticking up post-Claude Code | Yahoo Finance / Indeed Hiring Lab | 2026 | Indeed Hiring Lab data reported via Yahoo Finance showing a 15 percent rise in US software developer job postings since Claude Code's 2025 release, relevant to labour-market effects of AI-assisted coding. |
| f16 | AI cost overruns are adding up - with major implications for CIOs | CIO | 2025-10 | CIO.com reporting on enterprise AI budget misestimation, showing a majority of organizations misjudge AI project costs by more than 10 percent, relevant to TCO concerns in build-vs-buy decisions. |
| f17 | Nearly 7 in 10 firms report AI cost overruns | CFO Dive | 2026 | CFO Dive coverage of enterprise AI spending governance failures, indicating financial-side scrutiny of AI project ROI parallel to software delivery outcome concerns. |
| f18 | CIO, July 2026: AI is gaining momentum, cloud costs are rising, and cyber risks are on the increase | Brandsit (citing Wall Street Journal) | 2026-07 | References Wall Street Journal reporting on AI agent sprawl at companies including Lyft, DaVita, GitLab and FICO, illustrating enterprise governance gaps as agentic coding scales. |
| f19 | AI budgets soar, ROI still elusive | Computerworld | 2026-03 | Computerworld analysis of the pattern of promising AI pilots followed by cost overruns and organizational skepticism, directly relevant to whether AI investment is translating into measurable delivery outcomes. |
| f20 | Government Modernization Is Not a Software Problem | Cicero Institute | 2026-05 | Policy analysis citing the Reuters Pentagon investigation and OECD research on how high-accountability public-sector cultures distort innovation and IT project outcomes. |
| f21 | Developer Productivity Benchmarks 2026 | AI-Native Engineering Data | larridin.com | March 20, 2026 | Retrieved by this lane's web search. |
| f22 | Top 100 Developer Productivity Statistics with AI Tools 2026 | index.dev | Retrieved by this lane's web search. | |
| f23 | The AI Productivity Paradox Research Report | faros.ai | July 23, 2025 | Retrieved by this lane's web search. |
| f24 | 93% of Developers Use AI. Why Is Productivity Only 10%? | shiftmag.dev | March 31, 2026 | Retrieved by this lane's web search. |
| f25 | AI Coding Statistics - Adoption, Productivity & Market Metrics | getpanto.ai | June 7, 2026 | Retrieved by this lane's web search. |
Frontier Lab & Model News
Academic & arXiv
| ID | Title | Outlet | Date | Significance |
|---|---|---|---|---|
| a1 | The Empirical Reality of IT Project Cost Overruns: Discovering A Power-Law Distribution | SSRN / Journal of Management Information Systems | 2022-08 | Foundational peer-reviewed (JMIS) demonstration that IT project cost overruns are fat-tailed/power-law distributed across 5,392 projects, the key academic counter to normal-distribution assumptions in project risk management. |
| a2 | The Uniqueness of IT Cost Risk: A Cross-Group Comparison of 23 Project Types | Journal of Cost Analysis and Parametrics (SAGE) | 2026 | 2026 follow-up study testing whether IT cost overrun is more fat-tailed than 22 other project types, using 11,011 projects, extending Flyvbjerg et al.'s power-law findings. |
| a3 | Overspend? Late? Failure? What the Data Say About IT Project Risk in the Public Sector | arXiv | 2013 | arXiv working paper by Budzier and Flyvbjerg quantifying black-swan IT project risk and public-sector 'starved project' budget cuts, and critiquing Standish-style crisis narratives. |
| a4 | Why Your IT Project Might Be Riskier Than You Think | arXiv | 2013 | Companion arXiv paper (based on the Harvard Business Review 2011 article) quantifying the one-in-six black-swan IT project rate with 200% average cost overrun. |
| a5 | The Oxford Olympics Study 2016: Cost and Cost Overrun at the Games | arXiv | 2016 | Establishes Flyvbjerg's reference-class forecasting methodology later applied to IT project cost overrun research. |
| a6 | The Rise and Fall of the Chaos Report Figures | IEEE Software | 2009 | Canonical peer-reviewed critique arguing Standish CHAOS success/challenged definitions are methodologically meaningless for benchmarking. |
| a7 | The Rise and Fall of the Chaos Report Figures (ResearchGate) | ResearchGate | 2009 | Full-text version of the Eveleens and Verhoef critique showing political bias in IT forecasts undermines Standish benchmarking claims. |
| a8 | The Standish Report: Does It Really Describe a Software Crisis? | ResearchGate | 2006 | Academic paper assessing whether CHAOS-style figures represent a genuine software crisis or an artefact of biased sampling and definitions. |
| a9 | Software Process Models and Analysis on Failure of Software Development Projects | arXiv | 2013 | arXiv paper compiling multi-year Standish CHAOS trend data (1994-2009) alongside other failure-rate surveys such as TCS 2007. |
| a10 | Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity | arXiv / METR | 2025-07 | The core METR randomized controlled trial finding AI tools made experienced developers 19% slower despite believing they were faster, the most rigorous causal evidence on AI-assisted coding productivity. |
| a11 | Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (METR blog) | METR | 2025-07 | METR's own plain-language write-up including the July 2025 RCT design and a February 2026 update reporting a reversal to positive productivity with later AI tools. |
| a12 | We are Changing our Developer Productivity Experiment Design | METR | 2026-02 | METR's February 2026 methodological update describing a second RCT round (57 developers) and design changes addressing task-level versus developer-level randomization limitations. |
| a13 | Measuring AI Ability to Complete Long Tasks | METR | 2025-03 | METR's foundational time-horizon paper introducing the metric of AI task-completion length, showing a roughly 7-month doubling time, underpinning HCAST and RE-Bench based capability tracking. |
| a14 | How Does Time Horizon Vary Across Domains? | METR | 2025-07 | METR extension analysis testing whether the 7-month doubling trend generalizes beyond software/research tasks to other domains. |
| a15 | Time Horizon 1.1 | METR | 2026-01 | METR's 2026 update expanding the HCAST-derived task suite from 170 to 228 tasks, showing the evolving methodology behind autonomous capability benchmarking. |
| a16 | Clarifying limitations of time horizon | METR | 2026-01 | METR's own methodological caveats about benchmark selection bias and realism tradeoffs in its time-horizon capability metric. |
| a17 | IACDM: Interactive Adversarial Convergence Development Methodology | arXiv | 2026 | arXiv paper citing the METR RCT's three-way divergence between forecast, self-report, and measured productivity as motivation for a new AI-assisted development framework. |
| a18 | The Productivity-Reliability Paradox: Specification-Driven Governance for AI-Augmented Software Development | arXiv | 2026 | 2026 arXiv paper attempting to reconcile conflicting empirical findings from GitHub's RCT, GitClear's longitudinal churn analysis, and DORA's 2024 report on AI code quality versus delivery stability. |
| a19 | Adoption of low-code and no-code development: A systematic literature review and future research agenda | Journal of Systems and Software | 2024-11 | 2024 Journal of Systems and Software SLR of 40 primary studies (2017-2023) on LCNC and citizen development adoption, showing the academic evidence base is fragmented and largely qualitative. |
| a20 | What does current research say about the viability of low-code development? A systematic literature review | Journal of Systems and Software | 2026-04 | 2026 ScienceDirect SLR covering January 2021-January 2025 publications assessing LCNC viability, usability and production-readiness evidence. |
| a21 | An Empirical Study on Low-Code Programming using Traditional vs Large Language Model Support | arXiv | 2024-02 | arXiv empirical comparison of traditional low-code development against LLM-assisted low-code approaches. |
| a22 | The Evolution of Agile and Hybrid Project Management Methodologies: A Systematic Literature Review | arXiv | 2025-11 | November 2025 arXiv PRISMA-guided SLR tracing the shift from pure agile to hybrid frameworks and identifying implementation challenges and success factors over 8 years of literature. |
| a23 | Impact of Agile Methodology Use on Project Success in Organizations - A Systematic Literature Review | ResearchGate | 2020-12 | SLR (30 primary studies from 1507 papers, 2008-2019) explicitly noting the scarcity of large-scale empirical evidence for agile's causal effect on project success despite its widespread adoption. |
| a24 | A systematic literature review of agile software development projects | Journal of Systems and Software | 2025-03 | 2025 SLR of 208 articles (1999-2024) documenting persistent lack of convergence in agile software development research findings. |
| a25 | Reasons for the Failure of Information Technology Projects in the Public Sector | Springer | 2021 | Systematic literature review and meta-synthesis building a theoretical base for public-sector IT project failure factors, distinguishing governance and stakeholder drivers from private-sector patterns. |
VC & Analyst Reports
| ID | Title | Outlet | Date | Significance |
|---|---|---|---|---|
| v1 | The Rise and Fall of the Chaos Report Figures | IEEE Software / ResearchGate | 2010 | Foundational academic critique arguing Standish's success/challenged figures are not valid for benchmarking due to political bias in IT forecasts. |
| v2 | CHAOS Report on IT Project Outcomes | OpenCommons | 2025 | Aggregates Standish CHAOS figures across editions (1994-2020), showing the 31% success/50% challenged/19% failed 2020 baseline and government IT underperformance. |
| v3 | The Chaos Report - what makes projects successful | LinkedIn / Clear Star Group | 2025-02 | Practitioner summary of Standish's final 2020 CHAOS methodology (50,000 projects, 104 factors) and the finding that team quality dominates other success factors. |
| v4 | Why Your IT Project May Be Riskier than You Think | Harvard Business Review / ResearchGate | 2011 | Flyvbjerg and Budzier's early (HBR-linked) analysis of 1,471 IT projects establishing the fat-tail cost-overrun framing later expanded in Oxford Global Projects work. |
| v5 | Bent Flyvbjerg - University of Oxford profile and IT cost-risk research | Academia.edu / University of Oxford | 2009 | Documents Flyvbjerg's finding that IT project cost risk has a fatter tail than any of 22 other project types studied, the basis of the Oxford Global Projects thesis. |
| v6 | Software projects - the sting in the tail | Longevitas | 2023 | Applies Flyvbjerg & Gardner (2023) Appendix A data to show IT projects have a 73% mean cost overrun and a 447% conditional tail expectation for the worst 18% of projects. |
| v7 | Overspend? Late? Failure? What the Data Say About IT Project Risk in the Public Sector | arXiv | 2013 | Oxford academic paper by Budzier and Flyvbjerg specifically examining public-sector IT project risk data. |
| v8 | Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (coverage) | Independent practitioner analysis | 2025-07 | METR's randomised controlled trial finding AI tools made experienced developers 19% slower on complex tasks despite developers believing they were faster, the most rigorous independent AI-coding productivity study to date. |
| v9 | We are Changing our Developer Productivity Experiment Design | METR | 2026-02 | METR's own follow-up acknowledging its updated experiment shows unreliable signal due to selection bias, and that the original 20% slowdown estimate does not straightforwardly generalise. |
| v10 | What METR's Study Missed About AI Productivity in the Wild | Faros AI | 2026-02 | Industry telemetry (10,000+ developers) contrasting METR's lab findings with real-world organisational data showing task throughput up but delivery speed unchanged, illustrating the individual-versus-organisational productivity gap. |
| v11 | Announcing the 2025 DORA Report | Google Cloud Blog | 2025-09 | Google Cloud's DORA team confirms AI now correlates positively with software delivery throughput but continues to correlate negatively with delivery stability, reversing the 2024 finding. |
| v12 | DORA | Balancing AI tensions: Moving from AI adoption to effective SDLC use | DORA (Google Cloud) | 2026-03 | Articulates DORA's 'amplifier' thesis: AI magnifies the strengths of high-performing organisations and the dysfunctions of struggling ones rather than fixing underlying capability gaps. |
| v13 | DORA Report 2025 Key Takeaways: AI Impact on Dev Metrics | Faros AI | 2025-09 | Telemetry-based follow-up documenting an 'Acceleration Whiplash' pattern: rising throughput accompanied by sharply worsening PR review time, bug rates and incident rates in 2026 data. |
| v14 | State of AI-assisted Software Development 2025 | DORA / Google Cloud | 2025 | Primary DORA report PDF documenting the reversal from the 2024 finding that more AI use worsened delivery stability and throughput. |
| v15 | Investing in GitButler | Andreessen Horowitz | 2026-04 | a16z investment thesis describing developer role shifting from 'programmer' to 'system architect and agent orchestrator' as the basis for infrastructure bets in the coding-agent era. |
| v16 | Even a16z VCs say no one really knows what an AI agent is | TechCrunch | 2025-05 | Illustrates the definitional uncertainty within a16z's own infrastructure investing team about AI agents, useful for noting recency bias and hype-cycle dynamics in VC framing. |
| v17 | Delivering large-scale IT projects on time, on budget, and on value | McKinsey & Company | 2012 | McKinsey's canonical study with Oxford's BT Centre (5,400+ IT projects) establishing the 45% budget overrun, 7% schedule overrun, 56% value shortfall, and 17% black-swan-project baseline still cited across the industry. |
| v18 | PMI Talent Triangle: The Key to Project Management Success | PMI | 2025-03 | Summarises PMI Pulse of the Profession 2024 findings that project success rates are statistically similar across remote, hybrid and in-person delivery models. |
| v19 | Pulse of the Profession 2024: The Future of Project Work | Project Management Institute | 2024 | Primary PMI report finding no statistically significant difference in project performance between agile, hybrid and predictive methodologies, directly countering vendor-driven agile-superiority claims. |
| v20 | PMI's 2025 Pulse of the Profession report | Project Management Institute | 2025 | Introduces PMI's redefinition of project success as 'delivered value that was worth the effort and expense' rather than pure iron-triangle metrics. |
| v21 | Gartner Magic Quadrant for Low-Code Platforms for Enterprises (2026) | ToolJet | 2026-01 | Documents Gartner's current Leaders quadrant (OutSystems, Mendix, Power Apps, ServiceNow, Appian, Salesforce) for the low-code commissioning decision landscape. |
| v22 | Top 8 Low-Code, No-Code Platforms Reshaping Software Delivery in 2026 | GEM Corporation | 2026 | Cites Gartner's $44.5 billion 2026 low-code market forecast and Forrester's 87% enterprise developer adoption figure alongside vendor-specific lock-in details (e.g. Mendix migration requiring rebuilds). |
| v23 | Low-Code Hits $44.5B: Gartner 2026 Forecast Explained | byteiota | 2026-01 | Independent analysis warning of a coming 'technical debt reckoning' from ungoverned citizen-developer low-code adoption, a corrective to purely vendor-optimistic adoption-curve framing. |
| v24 | AI Copilot Code Quality: 2025 Data Suggests 4x Growth in Code Clones | GitClear | 2025-02 | GitClear's instrumented analysis of 211 million lines of code (2020-2024) showing refactoring collapse and duplicated-code growth as AI coding tools scaled, key independent code-quality evidence. |
| v25 | Report Summary: GitClear AI Code Quality Research 2025 | jonas.rs | 2025-02 | Detailed breakdown of GitClear churn statistics (5.5% to 7.9% two-week churn, refactoring collapse from 24.1% to 9.5%) used across the industry as independent evidence of AI-era code quality trends. |