Folder Explainer

Back to folder
GPT-5.6 Sol · Sync

Research Explainer · Flyvbjerg et al. (2026)

Most IT projects stay near budget, but a small minority run spectacularly over

Across 5,360 IT projects, 59% finished on or below budget. But the severe-overrun group averaged 453% above budget, creating a fat tail: rare outcomes so large that ordinary averages become dangerously reassuring.

Published February 2026

59.14% of IT projects finished on or below budget

18.26% ran more than 50% over budget

453% average overrun within that severe-overrun group

11,011 projects across 23 types in the full comparison

The result has two very different ends. Of 5,360 IT projects, 59.14% finished on or below budget. Yet 18.26% ran more than 50% over, and that severe-overrun group averaged 5.53 times its estimate, or 453% over budget.

The paper calls this a barbell pattern: a large group near budget at one end and a smaller group of disasters at the other. A fat tail is that second end stretched unusually far. Rare failures do not taper away quickly, so a handful can dominate the total. The frequency of overruns is not the main danger. Their potential size is.

Project typeProjects studiedAny overrunOver 50% above budgetAverage final cost in severe cases
IT5,36040.86%18.26%5.53× budget
Nuclear storage2592.00%52.00%5.27× budget
Roads2,08653.84%9.97%2.01× budget
Mining88649.44%16.93%2.29× budget
Rail37272.58%27.15%2.22× budget
Pipelines43755.38%9.61%2.05× budget
Read 5.53× budget as follows: the average severely over-budget IT project finished at 5.53 times its estimate, equivalent to a 453% overrun. Source: Flyvbjerg et al. (2026), Table 2.

How quickly extreme-overrun risk fades

How to read it: each bar is a tail-risk fade rate, called alpha in the paper. Shorter bars mean the chance of an even larger overrun falls more slowly. A score of 1 or below means extreme cases remain common enough that the fitted model has no stable long-run average. IT is the only project type beyond that threshold.

A log-scale chart showing the share of projects exceeding each cost-overrun size. IT's line stays above roads, rail, mining and pipelines and extends much further right, indicating much larger extreme overruns.
How to read it: the vertical axis is the share of projects whose overrun is at least this large. The horizontal axis is final cost divided by estimated cost, so 10 means ten times the budget. Both axes use compressed logarithmic scales. IT stays higher and extends much further right, showing that very large IT overruns remain more common long after the other project types have tapered away.

The paper gives each project type a tail-risk fade rate, technically the Pareto alpha. It asks a simple question: as overruns get larger, how quickly do they become rarer? Lower scores mean the danger fades more slowly. IT's central estimate was 0.92, the lowest of all 23 project types.

At a score of 1 or below, the fitted model has no stable long-run average or spread. No invoice is literally infinite. The point is that as more projects are observed, occasional disasters can keep dragging the calculated average upwards instead of letting it settle. An average-based contingency can therefore look precise while resting on a number that will not hold still.

IT's 95% confidence interval runs from 0.75 to 1.17, crossing the critical score of 1. So the evidence clearly shows an unusually heavy tail, but does not prove with certainty that the true population sits beyond that exact boundary. The risk is real; the mathematical category remains an estimate.

Risk levelTail-risk fade rate (paper: α)What it means in practiceProject types
Extremeα ≤ 1No stable average or spread; a few cases can dominate any sampleIT
Very high1 < α ≤ 2The average can settle, but volatility does not; contingency ranges remain unstableBuildings, dams, defence, fossil thermal power, hydroelectric dams, nuclear power, nuclear storage, Olympics, rail
High2 < α ≤ 3Average and spread can settle, but the results remain strongly lopsided and extreme-event proneAerospace, bridges, mining, pipelines, rail stations, roads, tunnels, water
Elevated3 < α ≤ 4Average, spread and lopsidedness settle; rare extremes still dominate tail-heavinessOil and gas
Moderateα > 4All four summary measures settle; conventional averages are more defensibleEnergy transmission, nuclear decommissioning, solar power, wind power
Skew describes how lopsided the outcomes are. Kurtosis describes how strongly rare extremes dominate the distribution. In the paper, an 'infinite' statistic means the fitted model has no stable long-run value for it, not that a project can cost infinity. Source: Flyvbjerg et al. (2026), Tables 3 and 4.

The dataset contains 11,011 projects across 126 countries and six continents, worth US$4.64 trillion in 2023 prices. The descriptive analysis used every project. Tail analysis began with the 5,596 projects whose actual cost exceeded the estimate, with the theoretical fits applied above selected tail cut-offs.

Cost overrun was actual divided by estimated real cost, with the estimate fixed at the final investment decision and actual cost measured at go-live. The authors compared Pareto and lognormal fits using quantile plots and three goodness-of-fit tests, then estimated alpha and its 95% confidence interval with a 500-iteration nonparametric bootstrap.

Sensitivity checks covered source, size, duration, completion year, classification and statistical method. No historical trend was detected, and neither duration nor estimated cost was significantly associated with overrun. Around 25% of candidate cost data was rejected for poor quality, while badly managed projects may be less likely to release records at all. The reported tail may therefore be the polite version.

Data sourceProjects
Owner accounts2,845
Freedom of information requests2,653
Public records4,141
Academic studies1,372
Total11,011
Flyvbjerg et al. (2026), Appendix B. Included projects by data source.

The paper begins with four explanations from prior research: immaturity, intangibility, goal ambiguity and stakeholder resistance. Each could amplify cost risk, and IT may combine all four more intensely than physical project types. The study does not measure them directly, so the data support the theory without proving its causal chain.

The cross-project pattern suggests two additions. The fattest-tailed types cluster around bespoke, tailor-made delivery, while modular solar, wind and transmission projects sit at the thin-tailed end. IT projects also averaged only 3.2 years against 6.9 years elsewhere, so short duration did not buy safety. The authors suggest that rushed, fast decisions may cancel the usual advantage of speed. That is a plausible diagnosis, not a causal result wearing a lab coat.

Four inherited explanations and two additions organise the paper's causal argument.

  • ImmaturityIT is a younger discipline with less accumulated standardisation, regulation and costing experience.
  • IntangibilitySoftware lacks physical milestones, remains unusually malleable and often hides integration trouble until work is under way.
  • Goal ambiguityUnclear objectives feed requirements volatility, scope creep and escalation.
  • Stakeholder resistanceIT changes workflows and power, giving affected users and outside actors reasons to obstruct implementation.
  • BespokenessTailor-made systems multiply novelty and interdependence, while modular components repeat what has already worked.
  • Fast thinkingShort cycles may compress planning and encourage biased decisions. Speed looks efficient until it starts feeding the tail.

This is an observational comparison, so it cannot establish final causality. The proposed mechanisms were not directly measured. The analysis covers upfront capital cost through go-live, not lifecycle cost, and it does not separate IT subtypes. Schedule evidence was preliminary, while benefit data was too sparse for analysis.

Project-type samples also vary sharply: nuclear decommissioning had 16 projects, the Olympics 21 and nuclear storage 25. Some distribution fits were inconclusive or produced conflicting tests, and the fitted IT Pareto tail used 48 observations above its selected cut-off. Its 95% confidence interval also crosses the alpha-equals-1 boundary. A serious tail does not grant immunity from sampling uncertainty.

For practice, the paper argues against normal-distribution assumptions and casual reliance on mean forecasts. It favours larger reference classes, more deliberate decisions, and standardised modular systems built from smaller repeatable parts. Those remedies still need causal testing. The paper has found the tail. It has not yet found the hand pulling it.

THE PRACTICAL CONSEQUENCE

Treat IT cost forecasts as tail-risk decisions, not ordinary estimates with a little contingency added. Use large reference classes, preserve extreme outcomes in the data, and reduce bespoke interdependence through modular design. The sensible response is not a more confident average. It is a project design that gives the tail fewer places to hide.

Reference

Flyvbjerg, B., Budzier, A., Aaen, J., Keil, M., & Zottoli, M. (2026). The uniqueness of IT cost risk: A cross-group comparison of 23 project types. Project Management Journal, 57(1), 14-43. https://doi.org/10.1177/87569728251340590

We use analytics cookies to understand site usage and improve the service. We do not use marketing cookies.