
AI & automation × Research note
LLMs on construction documents: what the evidence says about accuracy
Language models read contracts, specifications and incident reports fluently. Peer-reviewed research on the closest comparable field, legal documents, shows how often fluent is not the same as right, and which designs reduce the risk.
Key takeaways
- On verifiable questions about real court cases, general-purpose LLMs hallucinated 69% to 88% of the time in a peer-reviewed Stanford study.
- Commercial legal research tools built with retrieval still hallucinated 17% to 33% of the time. Retrieval reduces the problem; it does not remove it.
- The National Institute of Standards and Technology (NIST) lists confabulation as a core risk of generative AI.
- Accuracy improves sharply when the task is bounded, answers are grounded in cited source text, the system can decline, and a second method checks the first.
- Every AI tool on project documents needs an evaluation set drawn from your own documents before anyone relies on it.
01Why legal research is the right benchmark
Construction runs on documents: contracts and their clauses, specifications, RFIs, submittals, method statements, incident reports, minutes. The obvious AI use case is to ask a model questions about them. The obvious risk is that it answers confidently and wrongly.
There is little peer-reviewed research on LLM accuracy with construction documents specifically. The closest well-studied field is law: long, precise texts where a wrong answer has real consequences and where answers can be checked against a source. Two studies from Stanford’s RegLab are the most rigorous evidence so far.
02What the studies found
General-purpose models. Matthew Dahl and colleagues asked widely used LLMs verifiable questions about real court cases, questions with a checkable right answer such as who wrote an opinion or what a case held. Hallucination rates ranged from 69% for GPT-3.5 to 88% for Llama 2. Models also tended to accept false premises in the question, and were poor at judging their own confidence[1].
Specialist tools with retrieval. Vendors responded that retrieval-augmented generation (RAG), which fetches relevant source documents and asks the model to answer from them, solves the problem. Varun Magesh and colleagues ran the first preregistered test of leading legal AI research tools from LexisNexis and Thomson Reuters. They hallucinated between 17% and 33% of the time[2].
Figure
Share of answers containing a hallucination, by system type
Retrieval helps a lot, but a one-in-six to one-in-three error rate is not something you would accept from a quantity surveyor or a contracts manager. NIST’s Generative AI Profile of its AI Risk Management Framework names confabulation (confidently stated false content) as one of twelve core generative-AI risks[3].
Retrieval reduces hallucination. It does not remove it.
03Designs that reduce the error
The studies above tested open-ended questions. Accuracy is very different when the system is designed around a narrow, checkable task. Five design choices matter most:
- Bound the task. “Classify this incident narrative into one of nine causes” or “find the liquidated-damages clause” is far more reliable than “summarise the risks in this contract”.
- Ground every answer in source text. Show the clause, record or paragraph the answer came from, so a person can check it in seconds.
- Calculate with code, not with the model. Totals, counts and dates should come from a database query the user can see, not from generated text.
- Let it decline. A system that says “the data doesn’t answer that” is more useful than one that always answers.
- Use two readers. Where two independent methods agree, confidence is higher; where they disagree, send it to a person.
04Case study: Construction ERP Intelligence Copilot
From my portfolio · project 08 of 15
Synthetic data- Problem
- Managers need answers from ERP data without writing reports or queries.
- Decision supported
- Everyday questions on projects, costs, invoices and vendors.
- Data
- The ERP data sets from Projects 1, 6 and 7, with questions worded differently from training.
- Method
- Retrieval for record lookups, a router, eight parameterised SQL queries and guardrails that decline.
- Result
- 94.8% correct end to end: 100% of lookups, 85% of calculations, 100% of declines.
- Limits
- Synthetic data and template questions; real users phrase things more freely.
05What this looked like in my own tests
I have applied these principles in two projects, and the results show how much the design matters[4][5].
| Test | Data | Result |
|---|---|---|
| LLM coding injury causes, no training, definitions only | Real OSHA reports, 3,204 test cases | 81.4% vs 83.8% for a model trained on 14,000 reports |
| Trained model and LLM agree → auto-code | Same | 89.1% correct, covering 85% of reports |
| LLM extracting fall height from narratives | Same | Band matches OSHA in 98.5% of cases where it gives one |
| ERP chatbot: record lookups via retrieval | Synthetic ERP data, reworded questions | 100% correct |
| ERP chatbot: totals via parameterised SQL | Same | 85.0% correct |
| ERP chatbot: declining unanswerable questions | Same | 100% declined |
| ERP chatbot: end to end | Same | 94.8% correct |
The pattern is consistent with the research. Open-ended generation is where errors live. Classification against fixed definitions, extraction of a specific fact, retrieval of an exact record and calculation by query are where these tools become dependable, and where they can be tested properly. The whole OSHA run, 3,204 reports twice, cost under $1.20 at list price, so the barrier is evaluation effort, not model cost.
06How strong is the evidence?
Not every finding in this note rests on the same kind of evidence. This is how I would weigh each one before acting on it.
| Finding | Evidence | Strength | Main caveat |
|---|---|---|---|
| General-purpose LLMs hallucinate often on legal questions | Peer-reviewed study, large verifiable question set | Strong | 2023 models; newer models may differ |
| Retrieval reduces but does not remove hallucination | Preregistered study of commercial legal tools | Strong | Legal research, not construction documents |
| Bounded tasks with checks are far more accurate | My OSHA and ERP chatbot tests | Moderate | Two projects; ERP test on synthetic data |
| Confabulation is a core generative-AI risk | NIST AI 600-1 framework | Strong | A framework, not a measurement |
07Before you rely on an AI document tool
- Write down the task in one sentence. If you can’t, it is too open-ended to trust.
- Build an evaluation set from your own documents: 100–300 questions with answers checked by a person, including questions that should be declined.
- Measure against a simple baseline, such as keyword search or a checklist, not against nothing.
- Require citations to source text for every answer, and spot-check them.
- Keep calculations out of the model. Totals and dates come from queries the user can see.
- Route disagreement and low confidence to a person, and track the error rate monthly after go-live.
NotesSources
- Dahl, M., Magesh, V., Suzgun, M. and Ho, D. E. (2024). Large legal fictions: profiling legal hallucinations in large language models. Journal of Legal Analysis, 16(1), 64–93.
- Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D. and Ho, D. E. (2025). Hallucination-free? Assessing the reliability of leading AI legal research tools. Journal of Empirical Legal Studies.
- NIST (2024). Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1).
- Iwale, A. (2026). Construction Safety Intelligence: LLM versus trained model on OSHA reports. GitHub.
- Iwale, A. (2026). Construction ERP Intelligence Copilot: retrieval, text-to-SQL and guardrails evaluation. GitHub.
Figures are quoted from the sources above as published; where a source reports a range or a survey estimate, it is described that way. Results from my own projects say whether they use real public data or synthetic data.