hello@atuliwale.com
A worktable with floor-plan drawings, a scale ruler, a pencil, a tape measure and a white card model of a mid-rise building.
All insights

Estimating × Research note

Why construction estimates miss — and why a range beats a single number

Cost overruns are not bad luck. Decades of research show they are predictable, biased in one direction and correctable with a method that governments have used for twenty years: look at how similar projects actually turned out.

  • 5 min read
  • 3 sources
  • Atul Iwale
8.5%
of more than 16,000 large projects were delivered on budget and on time[1]
0.5%
were on budget, on time and delivered the benefits promised[1]
24–51%
government-recommended uplift to early cost estimates for standard vs non-standard buildings[2]

Key takeaways

  • Overruns are the norm: in Bent Flyvbjerg’s database of 16,000+ projects, 8.5% met both budget and schedule.
  • The errors are biased, not random. Early estimates are systematically too low, driven by optimism bias and, sometimes, strategic misrepresentation.
  • Reference class forecasting corrects the bias by starting from the actual outcomes of similar past projects, then adjusting.
  • Government appraisal guidance can build this in: one widely used version adds up to 24% to early capital cost for standard buildings and up to 51% for non-standard ones.
  • A single number hides the risk. A calibrated P10–P90 range, checked against what later happened, is more honest and more useful.

01Overruns are the norm, not the exception

Bent Flyvbjerg has spent three decades collecting cost and schedule data on large projects: buildings, rail, roads, bridges, IT, energy and more. In How Big Things Get Done (2023, with Dan Gardner), he reports that of more than 16,000 projects in the database, only 8.5% were delivered on budget and on time, and only 0.5% were on budget, on time and delivered the benefits promised[1].

Figure

Share of 16,000+ large projects that hit their targets

  • All projects in the database100%
  • On budget and on time8.5%
  • On budget, on time and on benefits0.5%
Share of more than 16,000 projects. Source: Flyvbjerg and Gardner (2023) [1].

Flyvbjerg calls this the “iron law”: over budget, over time, under benefits, over and over again. The important word is systematic. If estimating errors were random, some projects would come in well under and the average would be close to zero. They don’t. The distribution is skewed towards overruns, with a long tail of very large ones.

02Why early estimates are too low

The research points to two causes that usually act together[3]:

  • Optimism bias, the planning fallacy described by Daniel Kahneman and Amos Tversky. We build the estimate from the “inside view”: this project’s scope, this team’s plan, the risks we can see. We underweight the risks we can’t see yet, which on a building project include design development, ground conditions, approvals, scope growth and price movement.
  • Strategic misrepresentation. Where a low number helps a project get approved, there is pressure towards the low number. This is not always deliberate, but it pulls in the same direction as optimism.

Neither cause is fixed by working harder on the same estimate. A more detailed inside view is still an inside view. What corrects it is outside information.

A more detailed inside view is still an inside view. What corrects it is outside information.

03Reference class forecasting: start from what happened

Reference class forecasting, developed for projects by Flyvbjerg from Kahneman and Tversky’s work, takes the “outside view”[3]. It has three steps:

  1. Choose a reference class of past projects similar enough to be comparable: type, size, location, procurement route.
  2. Establish the distribution of their outcomes, for example actual cost against the estimate at the same stage.
  3. Place your project in that distribution and adjust only for differences you can actually justify.

Some governments build this into their appraisal rules. The best-known example, HM Treasury’s supplementary Green Book guidance, on optimism bias gives upper-bound uplifts to apply to capital cost at the earliest business-case stage, reducing as project-specific risks are identified and managed[2]:

Figure

HM Treasury optimism-bias upper bounds for capital expenditure

  • Standard buildings24%
  • Standard civil engineering44%
  • Non-standard buildings51%
  • Non-standard civil engineering66%
  • Equipment / development (incl. IT)200%
Upper bounds at outline business-case stage; lower bounds range from about 2% to 10%. Source: HM Treasury [2].

The point is not the exact percentages, which come from one government’s study of its public projects. It is the discipline: an early estimate should carry an explicit, evidence-based allowance, and that allowance should shrink only as risk is actually removed.

04From one number to a calibrated range

A single-point estimate says nothing about its own uncertainty. A range does, if it is honest. The usual way to express it is with percentiles:

  • P50: half of comparable outcomes came in below this figure, half above.
  • P10 to P90: the band that should contain about 80% of outcomes.

“Should” is the key word. A range is only useful if it is calibrated, meaning that when you check it against projects that later completed, about 80% of their actual costs fall inside the P10–P90 band. Too narrow and it gives false confidence; too wide and nobody can use it. Coverage on later projects is the test.

05Case study: Cost Plan Estimation (ML)

From my portfolio · project 03 of 15

Synthetic data
Problem
Feasibility-stage cost plans rely on coefficients that ignore how similar projects actually turned out.
Decision supported
Go/no-go and budget setting at feasibility, with an honest range.
Data
1,800 synthetic projects patterned on Indian cost plans, rebased to Jan-2018 prices.
Method
Tree models trained on 2018–2023 starts, tested on 2024–2025; quantile ranges.
Result
6.6% average error vs 11.2% for the coefficient method; P10–P90 covered 81% of later projects.
Limits
Synthetic data; must be re-tested on a firm’s own completed projects.

06How strong is the evidence?

Not every finding in this note rests on the same kind of evidence. This is how I would weigh each one before acting on it.

Strong: large official data sets or peer-reviewed studies. Moderate: a single study or a specific population. Indicative: surveys, vendor-backed reports or synthetic tests.
FindingEvidenceStrengthMain caveat
Most large projects overrun cost and scheduleDatabase of 16,000+ projects across sectorsStrongLarge projects; mix of sectors and countries
Early estimates are biased low, not randomly wrongConsistent finding across the overrun literatureStrongSize of bias differs by project type
Reference-class uplifts correct the biasGovernment appraisal guidance built on project dataModerateUplift values come from one country’s public projects
ML estimates with calibrated ranges beat coefficientsMy test on later projectsIndicativeSynthetic data; needs a firm’s own history

07What to change in your estimating workflow

  1. Keep an outcomes register. For every completed project, record the estimate at each stage and the final cost. This is your reference class.
  2. Rebase costs to a common date before comparing projects from different years.
  3. Start from the reference class, then adjust. Write down every adjustment and the evidence behind it.
  4. Report P10, P50 and P90, not just a figure, and say what drives the width.
  5. Test calibration yearly. Did about 80% of completed projects land inside your P10–P90 bands? If not, widen or narrow them.
  6. Test on later projects, not a random sample. Evaluate any estimating model on projects that started after its training data, as it will be used.

NotesSources

  1. Flyvbjerg, B. and Gardner, D. (2023). How Big Things Get Done. Currency / Penguin Random House.
  2. HM Treasury. Supplementary Green Book guidance: optimism bias.
  3. Flyvbjerg, B. (2006). From Nobel Prize to project management: getting risks right. Project Management Journal, 37(3), 5–15.

Figures are quoted from the sources above as published; where a source reports a range or a survey estimate, it is described that way. Results from my own projects say whether they use real public data or synthetic data.

Have a construction data problem worth solving?

Tell me about the process or decision you want to improve.

Let's talk →