
Contracts × Research note
Reading contract risk from the wording: what a neural network catches, and what it can’t
Contract value leaks through clauses nobody read closely enough, and construction disputes now run to tens of millions of dollars. A model that reads clause wording can put the riskiest clauses in front of a reviewer first, if it is built and tested with care.
Key takeaways
- World Commerce & Contracting estimates poor contracting erodes 8.6% of expected value on average, and 15% or more in complex industries.
- Arcadis’s disputes research reports major construction disputes averaging $43m and lasting 14.4 months; errors and omissions in contract documents are the top cause.
- Risk lives in word order: “within 30 days” versus “within 120 days”, “shall be limited” versus “shall not be limited”. Keyword lists miss it.
- On synthetic construction contracts, a 1D-CNN caught 95% of reviewer-rated high-risk clauses at 85% precision.
- Agreement with individual reviewers ranged from 85% to 89%, and one reviewer rated borderline clauses High more often: part of any model’s “error” is human disagreement.
01What poor contracting costs
World Commerce & Contracting (WorldCC), the international association for contracting professionals, has measured the gap between what contracts are expected to deliver and what they actually deliver for more than a decade. Its research puts average value erosion at 8.6%, often 15% or more in complex industries, down only slightly from 9.2% when it was first measured in 2014[1]. Construction is one of those complex industries.
When erosion turns into a dispute, the numbers get larger. Arcadis’s 2024 Construction Disputes Report found that the average dispute in North America was worth about $43 million and took about 14.4 months to resolve. The most common cause was errors and omissions in the contract document, followed by parties failing to understand or comply with their contractual obligations[2].
The most common cause of construction disputes is not the site. It is the contract document.
02Where the risk hides in a construction contract
The clauses that cause trouble are rarely hidden. They are simply long, numerous and similar enough that a tired reviewer reads what they expect to read. In construction and supply contracts, the recurring high-risk patterns include[3]:
- Unlimited liability: “shall not be limited” where the template said “shall be limited to”.
- Pay-when-paid: the subcontractor is paid only when the main contractor is, passing the client’s payment risk down the chain.
- Long payment terms: 90 or 120 days instead of 30, which quietly finances the other party.
- Auto-renewal with long notice periods: a contract that renews itself unless cancelled 120 days before expiry, reviewed 90 days before.
- Retention and escalation terms that differ from the commercial assumption in the estimate.
Notice that the risk is in the wording and word order, not in the presence of a keyword. “Limited” appears in both the safe and the dangerous version of a liability clause.
03Case study: Contract Lifecycle Intelligence
From my portfolio · project 04 of 15
Synthetic data- Problem
- High-risk clauses go unreviewed while contracts approach expiry and auto-renewal.
- Decision supported
- Which clauses and contracts need legal review first.
- Data
- 200 portfolio contracts (800 clauses) plus a library of 1,300 contracts and 6,539 reviewer-rated clauses.
- Method
- Rules engine for expiry and renewal; seven text models including a 1D-CNN that reads the wording.
- Result
- CNN catches 95% of High-risk clauses at 85% precision (keyword checklist macro-F1 0.38 vs 0.94).
- Limits
- Synthetic clauses are more regular than real ones; a screening aid, not legal advice.
04Teaching a model to read the wording
The project started as a rules engine: flag clauses whose risk category is already High, and flag auto-renewing contracts expiring within 90 days. That works, but it relies on someone having already rated every clause, which is the expensive part[3].
Version 2 rates clauses from their text. It learns from a library of 1,300 historical contracts and 6,539 clauses, each rated Low, Medium or High by one of six legal reviewers, and is tested on 800 later clauses it never saw. Seven methods were compared, from a keyword checklist to neural networks:
Figure
Macro-F1 on 800 unseen clauses (higher is better)
The convolutional network won because it detects short phrases in order, exactly where the risk lives. It caught 95% of High-risk clauses with 85% precision. Because a missed High clause costs far more than an extra review, High recall was tracked separately and the decision threshold was checked with a recall-weighted measure. At contract level, the model put 52 of 200 contracts on the priority list, including all 40 that contained a reviewer-rated High clause.
05Part of the error is human
One finding matters more than the headline score. When the models were cross-checked against individual reviewers, agreement ranged from 85% to 89%, and one reviewer rated borderline clauses High more often than the others[3]. Legal risk rating is a judgement, and judges differ.
That changes how to read any accuracy figure. A model cannot be more consistent than the labels it learns from, and some of what looks like model error is reviewers disagreeing with each other. In practice this argues for calibrating reviewers against each other on a shared sample before training, and for using the model as a screening and ordering aid, not a verdict.
06How strong is the evidence?
Not every finding in this note rests on the same kind of evidence. This is how I would weigh each one before acting on it.
| Finding | Evidence | Strength | Main caveat |
|---|---|---|---|
| Poor contracting erodes about 8.6% of value | WorldCC cross-industry research | Indicative | Survey-based estimate across industries |
| Contract document errors are the top dispute cause | Arcadis annual disputes survey | Moderate | North American cases reported to one firm |
| A CNN reading wording beats keyword rules | My test on 800 later clauses | Indicative | Synthetic clauses are more regular than real ones |
| Reviewer judgement caps model accuracy | Model agreement with six reviewers, 85–89% | Moderate | Measured on synthetic reviewer labels |
07Putting clause screening to work
- Decide what “high risk” means in writing, with examples, before anyone rates a clause.
- Calibrate reviewers on a shared sample and measure their agreement; that sets the ceiling for any model.
- Track recall on high-risk clauses separately. A missed unlimited-liability clause is not the same as a false alarm.
- Pull notice periods and payment terms into a register with dates, so renewals and deadlines are diarised, not remembered.
- Test on your own reviewed clauses before relying on any model; synthetic or vendor accuracy does not transfer.
- Keep a person in the loop. Use the model to order the review queue and highlight wording, not to approve contracts.
NotesSources
- World Commerce & Contracting. Poor contract management continues to cost companies 9% of their bottom line.
- Arcadis (2024). Disputes in the digital age: findings of the 2024 Construction Disputes Report.
- Iwale, A. (2026). Contract Lifecycle Intelligence, including the clause-risk NLP upgrade. GitHub.
Figures are quoted from the sources above as published; where a source reports a range or a survey estimate, it is described that way. Results from my own projects say whether they use real public data or synthetic data.