hello@atuliwale.com
A project engineer in a hi-vis vest at a site-office desk with a laptop, stacks of documents and rolled drawings, looking out at a building under construction.
All insights

AI & automation × Research note

LLMs on construction documents: what the evidence says about accuracy

Language models read contracts, specifications and incident reports fluently. Peer-reviewed research on the closest comparable field, legal documents, shows how often fluent is not the same as right, and which designs reduce the risk.

  • 5 min read
  • 5 sources
  • Atul Iwale
69–88%
hallucination rates of general-purpose LLMs on verifiable questions about real court cases[1]
17–33%
hallucination rates of commercial legal AI research tools built with retrieval (RAG)[2]
89%
accuracy where a trained model and an LLM agreed on OSHA injury coding (my analysis, real data)[4]

Key takeaways

  • On verifiable questions about real court cases, general-purpose LLMs hallucinated 69% to 88% of the time in a peer-reviewed Stanford study.
  • Commercial legal research tools built with retrieval still hallucinated 17% to 33% of the time. Retrieval reduces the problem; it does not remove it.
  • The National Institute of Standards and Technology (NIST) lists confabulation as a core risk of generative AI.
  • Accuracy improves sharply when the task is bounded, answers are grounded in cited source text, the system can decline, and a second method checks the first.
  • Every AI tool on project documents needs an evaluation set drawn from your own documents before anyone relies on it.

Construction runs on documents: contracts and their clauses, specifications, RFIs, submittals, method statements, incident reports, minutes. The obvious AI use case is to ask a model questions about them. The obvious risk is that it answers confidently and wrongly.

There is little peer-reviewed research on LLM accuracy with construction documents specifically. The closest well-studied field is law: long, precise texts where a wrong answer has real consequences and where answers can be checked against a source. Two studies from Stanford’s RegLab are the most rigorous evidence so far.

02What the studies found

General-purpose models. Matthew Dahl and colleagues asked widely used LLMs verifiable questions about real court cases, questions with a checkable right answer such as who wrote an opinion or what a case held. Hallucination rates ranged from 69% for GPT-3.5 to 88% for Llama 2. Models also tended to accept false premises in the question, and were poor at judging their own confidence[1].

Specialist tools with retrieval. Vendors responded that retrieval-augmented generation (RAG), which fetches relevant source documents and asks the model to answer from them, solves the problem. Varun Magesh and colleagues ran the first preregistered test of leading legal AI research tools from LexisNexis and Thomson Reuters. They hallucinated between 17% and 33% of the time[2].

Figure

Share of answers containing a hallucination, by system type

  • General-purpose LLM, best case (GPT-3.5)69%
  • General-purpose LLM, worst case (Llama 2)88%
  • Legal RAG tools, best case17%
  • Legal RAG tools, worst case33%
Sources: Dahl et al. (2024) [1]; Magesh et al. (2025) [2]. Different question sets, so compare the ranges rather than individual bars.

Retrieval helps a lot, but a one-in-six to one-in-three error rate is not something you would accept from a quantity surveyor or a contracts manager. NIST’s Generative AI Profile of its AI Risk Management Framework names confabulation (confidently stated false content) as one of twelve core generative-AI risks[3].

Retrieval reduces hallucination. It does not remove it.

03Designs that reduce the error

The studies above tested open-ended questions. Accuracy is very different when the system is designed around a narrow, checkable task. Five design choices matter most:

  1. Bound the task. “Classify this incident narrative into one of nine causes” or “find the liquidated-damages clause” is far more reliable than “summarise the risks in this contract”.
  2. Ground every answer in source text. Show the clause, record or paragraph the answer came from, so a person can check it in seconds.
  3. Calculate with code, not with the model. Totals, counts and dates should come from a database query the user can see, not from generated text.
  4. Let it decline. A system that says “the data doesn’t answer that” is more useful than one that always answers.
  5. Use two readers. Where two independent methods agree, confidence is higher; where they disagree, send it to a person.

04Case study: Construction ERP Intelligence Copilot

From my portfolio · project 08 of 15

Synthetic data
Problem
Managers need answers from ERP data without writing reports or queries.
Decision supported
Everyday questions on projects, costs, invoices and vendors.
Data
The ERP data sets from Projects 1, 6 and 7, with questions worded differently from training.
Method
Retrieval for record lookups, a router, eight parameterised SQL queries and guardrails that decline.
Result
94.8% correct end to end: 100% of lookups, 85% of calculations, 100% of declines.
Limits
Synthetic data and template questions; real users phrase things more freely.

05What this looked like in my own tests

I have applied these principles in two projects, and the results show how much the design matters[4][5].

Sources: [4][5]. The OSHA tests use real public data; the ERP tests use synthetic data and template questions, so real users will phrase things more freely.
TestDataResult
LLM coding injury causes, no training, definitions onlyReal OSHA reports, 3,204 test cases81.4% vs 83.8% for a model trained on 14,000 reports
Trained model and LLM agree → auto-codeSame89.1% correct, covering 85% of reports
LLM extracting fall height from narrativesSameBand matches OSHA in 98.5% of cases where it gives one
ERP chatbot: record lookups via retrievalSynthetic ERP data, reworded questions100% correct
ERP chatbot: totals via parameterised SQLSame85.0% correct
ERP chatbot: declining unanswerable questionsSame100% declined
ERP chatbot: end to endSame94.8% correct

The pattern is consistent with the research. Open-ended generation is where errors live. Classification against fixed definitions, extraction of a specific fact, retrieval of an exact record and calculation by query are where these tools become dependable, and where they can be tested properly. The whole OSHA run, 3,204 reports twice, cost under $1.20 at list price, so the barrier is evaluation effort, not model cost.

06How strong is the evidence?

Not every finding in this note rests on the same kind of evidence. This is how I would weigh each one before acting on it.

Strong: large official data sets or peer-reviewed studies. Moderate: a single study or a specific population. Indicative: surveys, vendor-backed reports or synthetic tests.
FindingEvidenceStrengthMain caveat
General-purpose LLMs hallucinate often on legal questionsPeer-reviewed study, large verifiable question setStrong2023 models; newer models may differ
Retrieval reduces but does not remove hallucinationPreregistered study of commercial legal toolsStrongLegal research, not construction documents
Bounded tasks with checks are far more accurateMy OSHA and ERP chatbot testsModerateTwo projects; ERP test on synthetic data
Confabulation is a core generative-AI riskNIST AI 600-1 frameworkStrongA framework, not a measurement

07Before you rely on an AI document tool

  1. Write down the task in one sentence. If you can’t, it is too open-ended to trust.
  2. Build an evaluation set from your own documents: 100–300 questions with answers checked by a person, including questions that should be declined.
  3. Measure against a simple baseline, such as keyword search or a checklist, not against nothing.
  4. Require citations to source text for every answer, and spot-check them.
  5. Keep calculations out of the model. Totals and dates come from queries the user can see.
  6. Route disagreement and low confidence to a person, and track the error rate monthly after go-live.

NotesSources

  1. Dahl, M., Magesh, V., Suzgun, M. and Ho, D. E. (2024). Large legal fictions: profiling legal hallucinations in large language models. Journal of Legal Analysis, 16(1), 64–93.
  2. Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D. and Ho, D. E. (2025). Hallucination-free? Assessing the reliability of leading AI legal research tools. Journal of Empirical Legal Studies.
  3. NIST (2024). Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1).
  4. Iwale, A. (2026). Construction Safety Intelligence: LLM versus trained model on OSHA reports. GitHub.
  5. Iwale, A. (2026). Construction ERP Intelligence Copilot: retrieval, text-to-SQL and guardrails evaluation. GitHub.

Figures are quoted from the sources above as published; where a source reports a range or a survey estimate, it is described that way. Results from my own projects say whether they use real public data or synthetic data.

Have a construction data problem worth solving?

Tell me about the process or decision you want to improve.

Let's talk →