
Production & quality × Research note
Rework and bad data: the cost line nobody budgets
Surveys and benchmarking studies agree that rework is a large, recurring cost, and that poor project information is one of its biggest causes. The fix starts long before the site: at the point where data is first entered.
- 14+ hrs
- a week per project team member spent on non-optimal activities: fixing mistakes, looking for data, resolving conflict[1]
- 48%
- of rework blamed on poor project data and miscommunication, in a survey of ~600 construction leaders[1]
- 2–20%
- of a project’s contract amount is typically lost to rework, per CII research[2]
Key takeaways
- In a survey of about 600 construction leaders, poor project data and miscommunication were blamed for 48% of rework.
- Team members reported spending 14+ hours a week on non-optimal activities, worth an estimated $177.5bn a year in labour cost in the market surveyed.
- Construction Industry Institute research puts rework at 2–20% of contract value and finds it can be predicted before construction starts.
- Bad data has recognisable shapes: duplicates, spelling variants, mixed units, codes from different standards and impossible dates. Every dataset in my portfolio had them.
- The cheapest place to fix data is where it is entered, with validation, ownership and one source of truth.
01What the surveys found
In 2018, FMI and PlanGrid surveyed nearly 600 construction leaders about how project teams spend their time and where things go wrong. Their report, Construction Disconnected, estimated that team members spend more than 14 hours a week on non-optimal activities such as looking for project data, fixing mistakes and managing conflict, worth about $177.5bn a year in labour cost across the market surveyed[1].
On rework specifically, respondents attributed 48% of it to poor project data and miscommunication: 26% to poor communication between team members and 22% to poor project information. In the market surveyed, that share represented an estimated $31.3bn of rework in a single year[1].
Figure
Share of rework attributed to each cause (survey estimate)
02What rework benchmarking says
The Construction Industry Institute (CII), a research consortium based at the University of Texas at Austin, has studied field rework for decades. Its research puts rework on a typical project at between 2% and 20% of the contract amount[2].
More usefully, CII developed the Field Rework Index, a short questionnaire completed before construction that predicts rework and cost growth. The factors with the strongest relationship to field rework are[2]:
- owner alignment,
- design rework,
- constructability commitment,
- interdisciplinary design coordination,
- the degree of project execution planning.
Projects with low index scores experienced low rework and even negative cost growth. Rework, in other words, is largely decided in design and planning, and it can be seen coming.
Rework is largely decided in design and planning, and it can be seen coming.
03Why bad data compounds
In an ERP-connected business, a record doesn’t stay in one place. A wrong cost code on a purchase order flows into commitments, then into the cost report, then into the forecast and the margin review. A supplier set up twice splits its spend and hides its delivery record. A drawing revision that doesn’t reach the site becomes rework.
That is why cleaning data at the reporting end is so expensive: every downstream report inherits the error, and every fix has to be repeated. It is far cheaper to stop the error where the data is first entered.
04Case study: Production Planning (RCC)
From my portfolio · project 15 of 15
Synthetic data- Problem
- Production orders with big output shortfalls and heavy rework need traceable review.
- Decision supported
- Which orders, items and processes to investigate.
- Data
- 700 production orders for 36 RCC items.
- Method
- Confirmed rule: shortfall of 15%+ and rework of 10%+; separate review of rework on orders that met plan.
- Result
- 70 confirmed orders on three items; 203 more orders had rework despite meeting or beating plan.
- Limits
- Thresholds chosen after seeing the data; a rule, not a model.
05What bad data actually looks like
“Poor project information” sounds abstract. In practice it has a small number of recognisable shapes. These are real examples from datasets in my own portfolio, two built on public government data and one on synthetic data designed to behave like real cost records[3]:
| Problem | Example found | What it breaks |
|---|---|---|
| Codes from different standards | 348 OSHA event codes stored at 2, 3 or 4 digits from more than one coding manual | Trend reports compare different things under one label |
| A rule change mid-series | OSHA’s January 2024 manual change moved the caught-in / struck-by line | A “trend” that is really a definition change |
| Duplicates | 128 duplicate NYC permit records; 6 duplicate rows in cost data | Double-counted cost, work or risk |
| Impossible dates | 21 NYC approvals dated before their filing | Negative durations, broken schedules |
| Mixed units | 18 cost entries recorded in sq ft instead of sq m | Rates wrong by a factor of about 10.8 |
| Spelling variants | Legacy city names (Bangalore, Bombay, Gurgaon) mixed with other spellings | One place split into several |
| Silent blanks | Missing soil or green-rating fields | Quietly dropped rows, biased averages |
None of these is exotic. Each one was either entered inconsistently or changed definition over time, and each one silently changes the answer to a management question unless someone catches it.
06Rework in an RCC production yard
Rework is easiest to see where production is measured order by order. My Production Planning (RCC) project reviews 700 production orders for 36 reinforced-concrete items, and separates two kinds of evidence[4]:
Figure
Production orders by rework pattern (700 orders)
The confirmed group is the obvious problem: large shortfalls with heavy rework, concentrated on just three items. The more interesting group is the 203 orders that met or beat their planned output and still had rework. On a production report they look fine. The rework only shows up if someone records it and looks for it, which is exactly why rework needs its own cost code and its own review.
07How strong is the evidence?
Not every finding in this note rests on the same kind of evidence. This is how I would weigh each one before acting on it.
| Finding | Evidence | Strength | Main caveat |
|---|---|---|---|
| Rework costs 2–20% of contract value | Construction Industry Institute research | Moderate | Wide range; depends on project type |
| Rework can be predicted before construction | CII Field Rework Index studies | Moderate | Mainly industrial projects |
| Poor data and communication cause ~half of rework | Survey of ~600 construction leaders | Indicative | Self-reported; vendor-sponsored report |
| Real data sets carry duplicates, unit and code errors | Counts from my own cleaning pipelines | Strong | Examples, not a rate across the industry |
08Controls that prevent it
- Validate at entry. Pick-lists for locations and categories, unit checks, date rules (approval cannot precede filing), mandatory fields that matter.
- Give master data an owner. Someone is accountable for the supplier, item and cost-code masters, and for merging duplicates.
- Version your definitions. When a coding standard changes, record the date, so analysis can separate the definition change from a real trend.
- One source for drawings and RFIs. Rework often starts when the site builds from a superseded revision.
- Measure rework explicitly. Give it a cost code. What isn’t measured is absorbed into “overrun” and never fixed.
- Score rework risk before construction. Use a structured check such as CII’s Field Rework Index at the end of design.
NotesSources
- FMI and PlanGrid (2018). Construction Disconnected: the high cost of poor data and miscommunication.
- Construction Industry Institute. The Field Rework Index: early warning for field rework and cost growth.
- Iwale, A. (2026). Portfolio project documentation (projects 13, 14 and 15: data-cleaning sections). GitHub.
- Iwale, A. (2026). Production Planning (RCC). GitHub.
Figures are quoted from the sources above as published; where a source reports a range or a survey estimate, it is described that way. Results from my own projects say whether they use real public data or synthetic data.