
Estimating × Research note
Why construction estimates miss — and why a range beats a single number
Cost overruns are not bad luck. Decades of research show they are predictable, biased in one direction and correctable with a method that governments have used for twenty years: look at how similar projects actually turned out.
Key takeaways
- Overruns are the norm: in Bent Flyvbjerg’s database of 16,000+ projects, 8.5% met both budget and schedule.
- The errors are biased, not random. Early estimates are systematically too low, driven by optimism bias and, sometimes, strategic misrepresentation.
- Reference class forecasting corrects the bias by starting from the actual outcomes of similar past projects, then adjusting.
- Government appraisal guidance can build this in: one widely used version adds up to 24% to early capital cost for standard buildings and up to 51% for non-standard ones.
- A single number hides the risk. A calibrated P10–P90 range, checked against what later happened, is more honest and more useful.
01Overruns are the norm, not the exception
Bent Flyvbjerg has spent three decades collecting cost and schedule data on large projects: buildings, rail, roads, bridges, IT, energy and more. In How Big Things Get Done (2023, with Dan Gardner), he reports that of more than 16,000 projects in the database, only 8.5% were delivered on budget and on time, and only 0.5% were on budget, on time and delivered the benefits promised[1].
Figure
Share of 16,000+ large projects that hit their targets
Flyvbjerg calls this the “iron law”: over budget, over time, under benefits, over and over again. The important word is systematic. If estimating errors were random, some projects would come in well under and the average would be close to zero. They don’t. The distribution is skewed towards overruns, with a long tail of very large ones.
02Why early estimates are too low
The research points to two causes that usually act together[3]:
- Optimism bias, the planning fallacy described by Daniel Kahneman and Amos Tversky. We build the estimate from the “inside view”: this project’s scope, this team’s plan, the risks we can see. We underweight the risks we can’t see yet, which on a building project include design development, ground conditions, approvals, scope growth and price movement.
- Strategic misrepresentation. Where a low number helps a project get approved, there is pressure towards the low number. This is not always deliberate, but it pulls in the same direction as optimism.
Neither cause is fixed by working harder on the same estimate. A more detailed inside view is still an inside view. What corrects it is outside information.
A more detailed inside view is still an inside view. What corrects it is outside information.
03Reference class forecasting: start from what happened
Reference class forecasting, developed for projects by Flyvbjerg from Kahneman and Tversky’s work, takes the “outside view”[3]. It has three steps:
- Choose a reference class of past projects similar enough to be comparable: type, size, location, procurement route.
- Establish the distribution of their outcomes, for example actual cost against the estimate at the same stage.
- Place your project in that distribution and adjust only for differences you can actually justify.
Some governments build this into their appraisal rules. The best-known example, HM Treasury’s supplementary Green Book guidance, on optimism bias gives upper-bound uplifts to apply to capital cost at the earliest business-case stage, reducing as project-specific risks are identified and managed[2]:
Figure
HM Treasury optimism-bias upper bounds for capital expenditure
The point is not the exact percentages, which come from one government’s study of its public projects. It is the discipline: an early estimate should carry an explicit, evidence-based allowance, and that allowance should shrink only as risk is actually removed.
04From one number to a calibrated range
A single-point estimate says nothing about its own uncertainty. A range does, if it is honest. The usual way to express it is with percentiles:
- P50: half of comparable outcomes came in below this figure, half above.
- P10 to P90: the band that should contain about 80% of outcomes.
“Should” is the key word. A range is only useful if it is calibrated, meaning that when you check it against projects that later completed, about 80% of their actual costs fall inside the P10–P90 band. Too narrow and it gives false confidence; too wide and nobody can use it. Coverage on later projects is the test.
05Case study: Cost Plan Estimation (ML)
From my portfolio · project 03 of 15
Synthetic data- Problem
- Feasibility-stage cost plans rely on coefficients that ignore how similar projects actually turned out.
- Decision supported
- Go/no-go and budget setting at feasibility, with an honest range.
- Data
- 1,800 synthetic projects patterned on Indian cost plans, rebased to Jan-2018 prices.
- Method
- Tree models trained on 2018–2023 starts, tested on 2024–2025; quantile ranges.
- Result
- 6.6% average error vs 11.2% for the coefficient method; P10–P90 covered 81% of later projects.
- Limits
- Synthetic data; must be re-tested on a firm’s own completed projects.
06How strong is the evidence?
Not every finding in this note rests on the same kind of evidence. This is how I would weigh each one before acting on it.
| Finding | Evidence | Strength | Main caveat |
|---|---|---|---|
| Most large projects overrun cost and schedule | Database of 16,000+ projects across sectors | Strong | Large projects; mix of sectors and countries |
| Early estimates are biased low, not randomly wrong | Consistent finding across the overrun literature | Strong | Size of bias differs by project type |
| Reference-class uplifts correct the bias | Government appraisal guidance built on project data | Moderate | Uplift values come from one country’s public projects |
| ML estimates with calibrated ranges beat coefficients | My test on later projects | Indicative | Synthetic data; needs a firm’s own history |
07What to change in your estimating workflow
- Keep an outcomes register. For every completed project, record the estimate at each stage and the final cost. This is your reference class.
- Rebase costs to a common date before comparing projects from different years.
- Start from the reference class, then adjust. Write down every adjustment and the evidence behind it.
- Report P10, P50 and P90, not just a figure, and say what drives the width.
- Test calibration yearly. Did about 80% of completed projects land inside your P10–P90 bands? If not, widen or narrow them.
- Test on later projects, not a random sample. Evaluate any estimating model on projects that started after its training data, as it will be used.
NotesSources
- Flyvbjerg, B. and Gardner, D. (2023). How Big Things Get Done. Currency / Penguin Random House.
- HM Treasury. Supplementary Green Book guidance: optimism bias.
- Flyvbjerg, B. (2006). From Nobel Prize to project management: getting risks right. Project Management Journal, 37(3), 5–15.
Figures are quoted from the sources above as published; where a source reports a range or a survey estimate, it is described that way. Results from my own projects say whether they use real public data or synthetic data.