hello@atuliwale.com
A thick contract open on a meeting table, with teal flags marking clauses, a fountain pen across the page and a closed laptop behind.
All insights

Contracts × Research note

Reading contract risk from the wording: what a neural network catches, and what it can’t

Contract value leaks through clauses nobody read closely enough, and construction disputes now run to tens of millions of dollars. A model that reads clause wording can put the riskiest clauses in front of a reviewer first, if it is built and tested with care.

  • 5 min read
  • 3 sources
  • Atul Iwale
8.6%
average value erosion from poor contracting, per World Commerce & Contracting[1]
14.4 mo
average length of a major construction dispute in Arcadis’s 2024 disputes report[2]
95%
of high-risk clauses caught by a CNN reading the wording (synthetic data), against a 0.38 macro-F1 keyword list[3]

Key takeaways

  • World Commerce & Contracting estimates poor contracting erodes 8.6% of expected value on average, and 15% or more in complex industries.
  • Arcadis’s disputes research reports major construction disputes averaging $43m and lasting 14.4 months; errors and omissions in contract documents are the top cause.
  • Risk lives in word order: “within 30 days” versus “within 120 days”, “shall be limited” versus “shall not be limited”. Keyword lists miss it.
  • On synthetic construction contracts, a 1D-CNN caught 95% of reviewer-rated high-risk clauses at 85% precision.
  • Agreement with individual reviewers ranged from 85% to 89%, and one reviewer rated borderline clauses High more often: part of any model’s “error” is human disagreement.

01What poor contracting costs

World Commerce & Contracting (WorldCC), the international association for contracting professionals, has measured the gap between what contracts are expected to deliver and what they actually deliver for more than a decade. Its research puts average value erosion at 8.6%, often 15% or more in complex industries, down only slightly from 9.2% when it was first measured in 2014[1]. Construction is one of those complex industries.

When erosion turns into a dispute, the numbers get larger. Arcadis’s 2024 Construction Disputes Report found that the average dispute in North America was worth about $43 million and took about 14.4 months to resolve. The most common cause was errors and omissions in the contract document, followed by parties failing to understand or comply with their contractual obligations[2].

The most common cause of construction disputes is not the site. It is the contract document.

02Where the risk hides in a construction contract

The clauses that cause trouble are rarely hidden. They are simply long, numerous and similar enough that a tired reviewer reads what they expect to read. In construction and supply contracts, the recurring high-risk patterns include[3]:

  • Unlimited liability: “shall not be limited” where the template said “shall be limited to”.
  • Pay-when-paid: the subcontractor is paid only when the main contractor is, passing the client’s payment risk down the chain.
  • Long payment terms: 90 or 120 days instead of 30, which quietly finances the other party.
  • Auto-renewal with long notice periods: a contract that renews itself unless cancelled 120 days before expiry, reviewed 90 days before.
  • Retention and escalation terms that differ from the commercial assumption in the estimate.

Notice that the risk is in the wording and word order, not in the presence of a keyword. “Limited” appears in both the safe and the dangerous version of a liability clause.

03Case study: Contract Lifecycle Intelligence

From my portfolio · project 04 of 15

Synthetic data
Problem
High-risk clauses go unreviewed while contracts approach expiry and auto-renewal.
Decision supported
Which clauses and contracts need legal review first.
Data
200 portfolio contracts (800 clauses) plus a library of 1,300 contracts and 6,539 reviewer-rated clauses.
Method
Rules engine for expiry and renewal; seven text models including a 1D-CNN that reads the wording.
Result
CNN catches 95% of High-risk clauses at 85% precision (keyword checklist macro-F1 0.38 vs 0.94).
Limits
Synthetic clauses are more regular than real ones; a screening aid, not legal advice.

04Teaching a model to read the wording

The project started as a rules engine: flag clauses whose risk category is already High, and flag auto-renewing contracts expiring within 90 days. That works, but it relies on someone having already rated every clause, which is the expensive part[3].

Version 2 rates clauses from their text. It learns from a library of 1,300 historical contracts and 6,539 clauses, each rated Low, Medium or High by one of six legal reviewers, and is tested on 800 later clauses it never saw. Seven methods were compared, from a keyword checklist to neural networks:

Figure

Macro-F1 on 800 unseen clauses (higher is better)

  • Keyword rules (baseline)0.38
  • Naive Bayes0.60
  • Linear SVM0.85
  • Logistic regression0.86
  • BiLSTM0.91
  • MLP neural network0.92
  • 1D-CNN (in the app)0.94
Synthetic clauses patterned on Indian construction contracts [3].

The convolutional network won because it detects short phrases in order, exactly where the risk lives. It caught 95% of High-risk clauses with 85% precision. Because a missed High clause costs far more than an extra review, High recall was tracked separately and the decision threshold was checked with a recall-weighted measure. At contract level, the model put 52 of 200 contracts on the priority list, including all 40 that contained a reviewer-rated High clause.

05Part of the error is human

One finding matters more than the headline score. When the models were cross-checked against individual reviewers, agreement ranged from 85% to 89%, and one reviewer rated borderline clauses High more often than the others[3]. Legal risk rating is a judgement, and judges differ.

That changes how to read any accuracy figure. A model cannot be more consistent than the labels it learns from, and some of what looks like model error is reviewers disagreeing with each other. In practice this argues for calibrating reviewers against each other on a shared sample before training, and for using the model as a screening and ordering aid, not a verdict.

06How strong is the evidence?

Not every finding in this note rests on the same kind of evidence. This is how I would weigh each one before acting on it.

Strong: large official data sets or peer-reviewed studies. Moderate: a single study or a specific population. Indicative: surveys, vendor-backed reports or synthetic tests.
FindingEvidenceStrengthMain caveat
Poor contracting erodes about 8.6% of valueWorldCC cross-industry researchIndicativeSurvey-based estimate across industries
Contract document errors are the top dispute causeArcadis annual disputes surveyModerateNorth American cases reported to one firm
A CNN reading wording beats keyword rulesMy test on 800 later clausesIndicativeSynthetic clauses are more regular than real ones
Reviewer judgement caps model accuracyModel agreement with six reviewers, 85–89%ModerateMeasured on synthetic reviewer labels

07Putting clause screening to work

  1. Decide what “high risk” means in writing, with examples, before anyone rates a clause.
  2. Calibrate reviewers on a shared sample and measure their agreement; that sets the ceiling for any model.
  3. Track recall on high-risk clauses separately. A missed unlimited-liability clause is not the same as a false alarm.
  4. Pull notice periods and payment terms into a register with dates, so renewals and deadlines are diarised, not remembered.
  5. Test on your own reviewed clauses before relying on any model; synthetic or vendor accuracy does not transfer.
  6. Keep a person in the loop. Use the model to order the review queue and highlight wording, not to approve contracts.

NotesSources

  1. World Commerce & Contracting. Poor contract management continues to cost companies 9% of their bottom line.
  2. Arcadis (2024). Disputes in the digital age: findings of the 2024 Construction Disputes Report.
  3. Iwale, A. (2026). Contract Lifecycle Intelligence, including the clause-risk NLP upgrade. GitHub.

Figures are quoted from the sources above as published; where a source reports a range or a survey estimate, it is described that way. Results from my own projects say whether they use real public data or synthetic data.

Have a construction data problem worth solving?

Tell me about the process or decision you want to improve.

Let's talk →