
Plant & equipment × Research note
Predicting equipment breakdowns: what the maintenance research promises, and what a model delivered
A breakdown on site costs the repair, the hire of a replacement and the crew standing idle. Predictive maintenance promises to catch failures early. Here is what the guidance claims, and what happened when I tested it on two years of telematics.
Key takeaways
- A widely used government maintenance guide puts the saving of a working predictive programme at 8–12% over preventive maintenance, and describes it as nearly eliminating catastrophic failures.
- Fixed alarm limits fire too late: in my test they caught only 16% of breakdowns in advance.
- Comparing each machine with its own normal is what gives the early warning, not absolute thresholds.
- At a realistic workshop capacity of 10 inspections a week, XGBoost caught 71% of breakdowns; at its cost-optimal threshold, 84% with a median of four weeks’ warning.
- Some failures give no warning. A model reduces breakdowns; it does not end them.
01Four ways to maintain a machine
One of the most widely used references is the Operations & Maintenance Best Practices guide from the Federal Energy Management Program (FEMP). It describes a ladder of maintenance strategies[1][2]:
| Strategy | When work happens | Trade-off |
|---|---|---|
| Reactive | After a failure | No planning cost, but the most downtime and collateral damage |
| Preventive | On a calendar or hour interval | Fewer failures, but parts are replaced whether or not they need it |
| Predictive | When condition data shows developing wear | Work only when needed, but needs sensors, data and analysis |
| Reliability-centred | A mix chosen per failure mode | Most effective, most analysis |
For predictive maintenance the guide is explicit about the prize: a properly functioning programme can save 8% to 12% over a preventive programme alone, and a well-run one will all but eliminate catastrophic equipment failures[1]. Those are claims about well-run programmes in facilities. The question for a contractor is whether the same signals work on mobile plant moving between dusty sites.
02Case study: Construction Assets (Fixed & Movable)
From my portfolio · project 12 of 15
Synthetic data- Problem
- Plant breaks down on site, and the failure pattern is only confirmed after it has happened.
- Decision supported
- Which machines the workshop should inspect this week.
- Data
- 80 telematics-fitted machines, 6,270 weekly readings and 249 breakdowns over two years.
- Method
- Seven classifiers on 33 features compared with each machine’s own normal; time-based test.
- Result
- XGBoost catches 71% of breakdowns with 10 inspections a week (OEM alarm limits: 16%).
- Limits
- Synthetic telemetry with known wear physics; savings depend on cost assumptions.
03The test: two years of plant telematics
My Construction Assets project started with a rule that confirms a failure pattern (three declining condition ratings with the same symptoms) only once it has fully happened. Useful for records, useless for Monday morning. The upgrade asks the plant manager’s real question: which machines are likely to break down in the next four weeks?[3]
- Fleet: 80 telematics-fitted machines, 6,270 weekly readings over two years (hours, load, vibration, temperature, hydraulic pressure, fault codes, monthly oil-iron samples) and 249 breakdowns.
- Features: 33, built only from past data, including each reading compared with the machine’s own normal level over the previous months, hours since service and weeks since the last repair.
- Test: trained to December 2025, tuned on January–March 2026, tested once on April–August 2026.
- Baselines: the manufacturer’s alarm limits and a service-overdue rule.
04What the models caught
Figure
Breakdowns caught in advance with 10 inspections a week (test months)
Fixed alarm limits did worst because they fire when the reading is already abnormal in absolute terms, by which time the failure is close. The models watch for a machine drifting away from its own normal, weeks earlier.
In cost terms, at its cost-optimal threshold XGBoost caught 69 of 82 test-period breakdowns (84%) with a median of four weeks’ warning, cutting breakdown-related cost in the simulation from Rs 422 lakh to Rs 212 lakh. By cause, it caught all bearing and gear wear and 95% of overheating, but far fewer sudden failures, which by nature give little warning[3].
The warning comes from a machine drifting away from its own normal, not from crossing a fixed limit.
05Why trees beat the SVM here
- Messy data favours tree models. A third of the fleet has no hydraulic sensor, oil results arrive monthly and telematics drops out. XGBoost learns what a missing value means; an SVM needs values imputed, and an imputed median looks like a real reading.
- Context interacts. A temperature rise means different things for an electric hoist and a diesel excavator in a hot summer. Trees split on both; a single distance measure treats them alike.
- Ranking beats thresholds. A workshop has fixed capacity, so “inspect the top ten each week” is the realistic operating rule, and it is where the gap between models shows most clearly.
- A second opinion helps. Where XGBoost and the SVM disagree strongly, that machine deserves a closer look.
06How strong is the evidence?
Not every finding in this note rests on the same kind of evidence. This is how I would weigh each one before acting on it.
| Finding | Evidence | Strength | Main caveat |
|---|---|---|---|
| Predictive maintenance saves 8–12% over preventive | Government (FEMP) maintenance guidance | Moderate | Facility equipment, not mobile plant |
| Fixed alarm limits warn too late | My test against OEM-style limits | Indicative | Synthetic telematics |
| Tree models suit patchy sensor data | Seven-model comparison, time-split test | Indicative | Five-month synthetic test window |
07Starting predictive maintenance on a fleet
- Record every breakdown properly: date, machine, root cause, downtime and cost. Without outcomes there is nothing to learn from.
- Collect condition data you already have (hour meters, fault codes, oil samples) before buying new sensors.
- Compare each machine with its own history, not only with fixed limits.
- Plan to workshop capacity: rank machines and inspect the top N each week.
- Price both errors: the cost of an unnecessary inspection against the cost of a missed breakdown, including hire and idle crews.
- Test on later months than you trained on, and retrain as the fleet and sites change.
NotesSources
- US Department of Energy, Federal Energy Management Program (2010). Operations & Maintenance Best Practices: A Guide to Achieving Operational Efficiency, Release 3.0.
- Pacific Northwest National Laboratory. O&M best practice issue discussion: maintenance approaches.
- Iwale, A. (2026). Construction Assets, including the predictive-maintenance upgrade. GitHub.
Figures are quoted from the sources above as published; where a source reports a range or a survey estimate, it is described that way. Results from my own projects say whether they use real public data or synthetic data.