reportAI HallucinationsVendor Watch
How often does legal AI hallucinate? The measurement register, September 2026
Every published measurement of legal-AI error on one schema, set against the vendor claims: 17% to 33% for research tools, 43% to 88% for general models.
Edited and verified by Cognesio LLP · updated
Researched with AI assistance · sources verified by Cognesio LLP · How this was made ↓
This report collects every published measurement of how often AI fabricates or misstates the law, in the United States and elsewhere, on one schema: who measured, what task, how many queries, what counted as a hallucination, which tool and version, when, and the rate. It then sets those measurements against what the vendors say about their own products. It is the second edition of a page first published on 16 July 2026, expanded from two studies to a register of nine entries and four vendor claims, and it will be updated as new measurements appear.
The short answer has not changed since 2024, because nobody has re-run the only rigorous test of the legal-research tools. On 202 preregistered legal queries in spring 2024, Lexis+ AI hallucinated in 17 percent of responses and answered 65 percent accurately; Westlaw’s AI-Assisted Research hallucinated in 33 percent and answered 42 percent accurately; GPT-4 hallucinated in 43 percent. Vendor-participating benchmarks since then report task accuracy of 54 to 95 percent, with different tasks and different graders, and the incident record shows 150 court decisions in which fabricated authority reached a court anyway, 93 of them dated 2026, to 3 September.
Key findings
- Nine entries, eight measurements and one theoretical account, have been published between January 2024 and August 2026, by four research groups, one benchmark company and two vendors. They use at least four different definitions of “hallucination” and are not comparable with each other.
- The only measurement of the major legal-research tools by researchers with no commercial interest (Magesh, Surani, Dahl, Suzgun, Manning and Ho, Journal of Empirical Legal Studies, 2025; queries run 22 March to 27 May 2024) found hallucination in 17 percent of Lexis+ AI responses, 17 percent of Ask Practical Law AI responses, 33 percent of Westlaw AI-Assisted Research responses and 43 percent of GPT-4 responses.
- Thomson Reuters answered that study on 10 June 2024 with an internal accuracy figure of “approximately 90%” for AI-Assisted Research, graded by two lawyers per question; the methods and the questions were not published, and the two numbers measure different things.
- LexisNexis’s launch claim of 25 October 2023, “hallucination-free linked legal citations”, is a claim about citation links, not about the correctness of the answer; the 2025 study found 17 percent of Lexis+ AI responses hallucinated under its definition, which includes correct citations that do not support the proposition.
- Two Vals AI benchmarks with vendor participation (27 February 2025 and 23 October 2025) report task accuracy of 54 to 95 percent for participating tools against lawyer baselines of 50 to 80 percent; Thomson Reuters and LexisNexis declined to enter the legal-research benchmark and vLex withdrew before the results were published, and the February grading was partly automated.
- General-purpose models measured on verifiable questions about US federal cases (Dahl, Magesh, Suzgun and Ho, Journal of Legal Analysis, 2024) hallucinated between 58 percent (GPT-4) and 88 percent (Llama 2) of the time.
- Two benchmarks and one position paper published in May and June 2026 address adjacent questions, not the rate itself: whether models can detect fabricated citations in a brief (GPT-5 reached 84.4 percent recall but 55.0 percent F1), whether retrieval-augmented legal systems can be evaluated at the claim level (the abstract motivates the benchmark on prior work finding that they still hallucinate at varying rates), and whether benchmark scores survive contact with a litigant’s own words (the authors argue current benchmarks cannot show that they do).
- The incident tracker’s 150 rows are a lagging indicator, not a rate: 92 of the 150 orders do not name the tool, and the rows dated 2026 (to 3 September) are 93 against 40 in 2025 because sweeps deepened, not because any measured rate rose.
Why this question
“How often does legal AI hallucinate” is the question behind every court rule on verification, every bar guidance note and every sanctions order in the incident tracker. It is also the question vendors answer with a marketing number and researchers answer with a benchmark number, and the two never describe the same thing. A general counsel deciding whether to allow a tool, a judge deciding whether to require certification, and a lawyer deciding how much to check all need the same three facts: what was measured, on what, and by whom with what interest.
The search record shows the demand: “ai with lowest hallucination rate” and “ai hallucination rate” carry 390 and 140 searches a month, and the pages that answer them are aggregators of vendor benchmarks. Nobody had put the measurements and the claims in one table with their methods exposed.
Method and data
A measurement is included if it reports a rate of hallucination, error or accuracy for a named AI system on a legal task, with a published method. Vendor blog posts count as claims, not measurements, unless the method is published. Each row of the register records: source and date; the systems tested; the task; the number of queries; who graded and how; the definition of hallucination or error; the reported rate. The register was compiled from the primary documents listed in Appendix B, each read for this edition (the ClaimRAG-LAW entry from its abstract only), and from the tools directory, which records each vendor’s published accuracy evidence with the field accuracyEvidence.
The incident data is the site’s incident collection, 150 records as of 5 September 2026, used only for the counts stated. Rates cannot be derived from it: the tracker records decisions that reached a court and were found, not the denominator of uses.
Disclosure: SafeLegalAI is published by Cognesio LLP, whose team also builds LegalAI Space, a governance product listed in the tools directory with an ownership label. LegalAI Space publishes no accuracy claim and is not in the register.
The register
| # | Measurement (date) | Systems | Task and size | Grading | Definition | Reported rate |
|---|---|---|---|---|---|---|
| 1 | Dahl, Magesh, Suzgun, Ho, Large Legal Fictions, arXiv 2 Jan 2024, J. Legal Analysis 2024 | GPT-4, GPT-3.5, PaLM 2, Llama 2 | Fourteen query types about a random sample of US federal cases (existence, citation, holding, authorship and others) | Reference-based: answers checked against case metadata | ”Textual output that is not consistent with legal facts” | Hallucination in 58% (GPT-4) to 88% (Llama 2) of responses on reference-based tasks; GPT-3.5 69%, PaLM 2 72% |
| 2 | Magesh, Surani, Dahl, Suzgun, Manning, Ho, Hallucination-Free?, arXiv 30 May 2024, J. Empirical Legal Studies 2025 | Lexis+ AI, Westlaw AI-Assisted Research, Ask Practical Law AI, GPT-4 | 202 preregistered open-ended legal research queries in four categories; run 22 Mar to 22 Apr 2024 (Lexis, Practical Law, GPT-4) and 23 to 27 May 2024 (Westlaw) | Manual, two axes: correctness and groundedness; inter-rater validation | Incorrect (misstates the law or a fact) or misgrounded (correct statement, cited source does not support it) | Hallucinated: Lexis+ AI 17%, Ask Practical Law AI 17%, Westlaw AI-AR 33%, GPT-4 43%. Accurate (correct and grounded): Lexis+ AI 65%, Westlaw 42%, Practical Law 20%, GPT-4 49%. Incomplete: Practical Law 63%, Westlaw 25%, Lexis 18%, GPT-4 8% |
| 3 | Thomson Reuters, statement by Mike Dahn, Head of Westlaw Product Management, 10 Jun 2024 | Westlaw Precision AI-Assisted Research | ”Hundreds of real-world legal research questions”; internal | Two lawyers per result, senior lawyer arbitration | Not published | ”Accuracy rate of approximately 90%” (internal; method and questions unpublished) |
| 4 | Harvey, Introducing BigLaw Bench, 29 Aug 2024 | Harvey assistant models, foundation models | Real-world legal work tasks (drafting, analysis); vendor-built | Answer score (requirements met minus errors incl. hallucinations); source score (claims supported by an accurate source) | Errors including hallucinated sources (document text or page number) | Harvey outputs “74% of a final, expert lawyer-quality work product”; foundation models “consistently hallucinate sources” (vendor-run) |
| 5 | Vals AI, Vals Legal AI Report, 27 Feb 2025 | CoCounsel (Thomson Reuters), Vincent AI (vLex), Harvey Assistant, Oliver (Vecflow); Lexis+ AI withdrew from the sections reported | Seven tasks (data extraction, document Q&A, summarisation, redlining, transcript analysis, chronology, EDGAR research); 240 questions across the seven tasks, over 500 graded samples | LLM-as-judge with lawyer baseline sourced by Cognia Law | Accuracy per task, not hallucination | Task scores 53.6% to 94.8% (lowest: Vincent AI on redlining; highest: Harvey on document Q&A); lawyer baseline 50.3% to 80.2% by task; Harvey 94.8% on document Q&A; CoCounsel averaged 79.5% across four tasks; Vincent 53.6% to 72.7% and Oliver 55.2% to 74.0% across the tasks they entered |
| 6 | Kalai, Nachum, Vempala, Zhang, Why Language Models Hallucinate, arXiv 4 Sep 2025 | General | Theoretical, with evaluation-incentive analysis | Not applicable | Statistical errors rewarded by evaluations that penalise abstention | No rate; explains why rates persist: “training and evaluation procedures reward guessing over acknowledging uncertainty” |
| 7 | Vals AI, VLAIR Legal Research, 23 Oct 2025 | Alexi, Counsel Stack, Midpage, ChatGPT (GPT-4) | 210 legal-research questions across nine research types | Blind grading by a consortium of law firms and academics | Accuracy of the answer | Lawyer baseline 71%; Alexi 80%, Counsel Stack 81%, Midpage 79%, ChatGPT 80%; Thomson Reuters, LexisNexis and vLex did not participate |
| 8 | Das, Abualhaija, Bianculli, ClaimRAG-LAW, arXiv 20 May 2026 (rev. 14 Aug 2026) | State-of-the-art legal and general RAG systems (unnamed in the abstract) | Claim-level evaluation of retrieval and generation, French and English | Claim-level automated analysis | Claim-level support of generated responses | No rate in the abstract; the paper motivates itself on prior work showing that “RAG systems, whether general-purpose or legal-specific, still hallucinate at varying rates” and reports fine-grained retrieval and generation limitations (rates in the paper body, not read for this edition) |
| 9 | Liu, Stammbach, Henderson, Who Checks the Citations?, arXiv 19 Jun 2026 (rev. 6 Aug 2026) | Five models in agentic and non-agentic settings, plus Claude Code; best result from GPT-5 | Detecting injected fabricated citations in 1,300 brief excerpts | Automated against injected errors | Detection of a fabricated citation, not generation of one | GPT-5 84.4% recall, 55.0% F1 in the agentic setting, 15.3 steps per excerpt; “all models struggle with subtle error categories” |
A tenth entry belongs in the register by omission: as of 5 September 2026 no independent re-test of Lexis+ AI, Westlaw AI-Assisted Research or CoCounsel on open-ended legal research had been published in the venues searched for this edition (arXiv, the Stanford RegLab publications page, Vals AI’s reports and the Journal of Empirical Legal Studies). Every number a buyer can cite for those products in September 2026 is either the 2024 study, a vendor’s internal figure, or a benchmark the vendor entered or declined.
Four definitions, four numbers
The register’s rates cannot be averaged because “hallucination” means four different things in it.
Not consistent with legal facts (row 1). The 2024 study asked models verifiable questions about real cases and scored the answer against the record. A wrong author, a wrong year or a non-existent case all count. The 58 to 88 percent range is a rate of false statements about cases the model was asked about directly.
Incorrect or misgrounded (row 2). The 2025 study added the second axis: a correct statement of law supported by a citation to a source that does not say it is a hallucination. This is the definition that matters for a brief, because a court checks the citation, not the proposition. It is also why the study’s 17 percent for Lexis+ AI is compatible with LexisNexis’s statement that its citations are linked and Shepardized: a real, valid, linked case can still be cited for a proposition it does not contain. The study gives Table 4 examples of Lexis+ AI citing the overruling case as authority for the overruled rule.
Task accuracy (rows 3, 5, 7). The vendor and benchmark figures score whether the answer to a task was right, as graded by a lawyer, two lawyers or a model. A response can be accurate on this definition and still contain a misgrounded citation; a response can be inaccurate without hallucinating anything. Thomson Reuters’ 90 percent and the Stanford 42 percent for the same product are not in conflict on their face, because one counts accurate answers to the vendor’s questions and the other counts hallucinated responses to preregistered questions the vendor did not choose.
Fabricated source (rows 4 and 9). BigLaw Bench’s source score and the June 2026 detection benchmark both treat the citation as the unit: does it exist, does it say that, can a model tell. This is the definition closest to what the incident tracker records, and it is the one on which the 2026 research says even GPT-5, working agentically over fifteen steps, misses a substantial share of subtle fabrications.
Vendor claims against the evidence
| Vendor and product | Claim (date, source) | What the claim covers | Nearest measurement | What the measurement found |
|---|---|---|---|---|
| LexisNexis, Lexis+ AI | ”hallucination-free linked legal citations” (press release, 25 Oct 2023); “minimizes the risk of invented content, or hallucinations, and checks all citations against Shepard’s” | That linked citations resolve to real, validated documents | Row 2 (Mar to Apr 2024) | 17% of responses hallucinated (incorrect or misgrounded); 65% accurate; the highest accuracy of the tools tested. Lexis+ AI withdrew from the sections of the Feb 2025 Vals report |
| Thomson Reuters, Westlaw AI-Assisted Research | ”accuracy rate of approximately 90%” in internal testing; the product “can occasionally produce inaccuracies” (statement of 10 Jun 2024) | Accuracy on the vendor’s own question set, two-lawyer graded | Row 2 (May 2024) | 33% of responses hallucinated; 42% accurate; “nearly twice as often as the other legal tools” |
| Thomson Reuters (Casetext), CoCounsel | CoCounsel “does not make up facts, or ‘hallucinate’” (Casetext, 2023, quoted in Magesh et al., n.2) | Grounding in known data sources | Row 5 (Feb 2025) | Averaged 79.5% task accuracy across the four tasks entered in the Vals report; no independent hallucination measurement published |
| Harvey | Assistant models produce “74% of a final, expert lawyer-quality work product”; foundation models “consistently hallucinate sources” (BigLaw Bench, 29 Aug 2024) | Vendor-built benchmark of the vendor’s own product | Row 5 (Feb 2025) | Highest single task score in the Vals report (94.8%, document Q&A); tied the lawyer baseline on chronology; no independent hallucination rate published |
Two patterns. First, every vendor claim is phrased about grounding or citation validity, and every independent measurement that found a problem found it in the proposition attached to a valid citation. Second, the two largest research vendors have each published one number since 2023, both internal, and neither has entered an independent benchmark of open-ended legal research: Thomson Reuters and LexisNexis did not enter the October 2025 Vals research benchmark, vLex withdrew from it, and Lexis+ AI withdrew from the reported sections of the February 2025 report. The tools directory records the published accuracy evidence per vendor: of 47 verified independent tools, 18 publish any accuracy evidence; 14 publish vendor evidence, 6 publish third-party evidence and 2 publish both (24 evidence entries in all, 15 vendor and 9 third-party).
The incident record as a lagging indicator
The tracker holds 150 decisions as of 5 September 2026 in which fabricated or misstated authority reached a court. Fifty-eight of the 150 orders name a tool. Counting rows that may name more than one: ChatGPT in 27, GPT-4o in 1, Claude in 5, Google products in 4, Westlaw or CoCounsel in 3 (one of them a row in which counsel denied using ChatGPT and named Westlaw’s tools), Microsoft Copilot in 2, a practice-management or specialist tool in 7 (LEAP, Visto.ai, Legal Genius, Strongsuit, ChatOn, a firm’s internal tool, Open Law), and 14 that say only that generative AI was used or suspected. Ninety-two do not identify the product.
Those counts are not rates. The denominator (how many filings used each tool) is unknown; the numerator is what courts wrote down and what the tracker found. The tool named most is the tool most people use. The three rows naming Westlaw or CoCounsel, one of them the Lacey v State Farm row naming CoCounsel, Westlaw Precision and Gemini together, are the only points where the register’s products and the incident record touch, and in each the court’s finding was that the lawyer did not check, not that the tool’s rate was known.
What the record does measure is that the mechanism the 2025 study described (a real citation attached to a proposition it does not support) is the one courts keep finding. The study’s typology of misgrounded responses reads as a catalogue of the tracker: holdings reversed, litigant arguments attributed to the court, overruled cases stated as good law without the citation. The rate at which that happens in practice remains unmeasured; the rate at which it is caught and sanctioned is in the sanctions ledger.
What a 2026 re-test would need
The 2025 study’s method supplies most of them, and its limitations section adds the hardest one: a preregistered query set, a definition that scores groundedness as well as correctness, manual grading with inter-rater checks, product versions and dates recorded, and access, because the vendors restrict their interfaces and the sample stayed at 202 queries. The June 2026 detection benchmark adds one: the graders themselves are now models, and the best of them misses fabrications that a human cite-checker would catch. A re-test that used a model to grade would inherit the error it was measuring.
Until one is published, the honest summary for a reader in September 2026 is the one this page carried in July: the best legal-research tool tested hallucinated in about one response in six on preregistered open-ended queries in spring 2024; the general chatbots most litigants use hallucinated in the majority of verifiable case questions in late 2023; the vendors report higher accuracy on their own tasks by their own graders; and no court in the tracker has treated any of those numbers as a reason not to check.
What to watch
The Stanford RegLab group has not announced a re-test; if one appears it will supersede row 2. Vals AI publishes new task reports through the year, and the participation list is the first thing to read in each. The detection benchmark (row 9) is the likeliest to be re-run against new model releases, and its recall figure is the number that bears on whether AI cite-checking can be trusted. On the incident side, the tracker’s tool-named share (58 of 150) will move as courts start requiring the tool to be identified under rules like Montana’s Fourth District Rule 3.G and Judge Graham’s standing order.
Three sentences journalists can quote
The only independent measurement of the major legal-research tools, on 202 preregistered queries in spring 2024, found that Lexis+ AI hallucinated in 17 percent of responses, Westlaw AI-Assisted Research in 33 percent and GPT-4 in 43 percent; it has not been repeated.
Thomson Reuters’ response was an internal accuracy figure of approximately 90 percent for the same product, graded by two lawyers on questions it has not published.
In the SafeLegalAI incident tracker, 92 of 150 court decisions on fabricated authority do not name the tool, so the record shows which products were admitted to, not which ones fail most.
Appendix A: the register as data
| # | Source | Date | Independent of vendors | Rate type | Headline figure |
|---|---|---|---|---|---|
| 1 | Dahl et al., J. Legal Analysis | Jan 2024 (arXiv), 2024 (journal) | Yes | Hallucination, reference-based | 58% to 88% |
| 2 | Magesh et al., J. Empirical Legal Studies | May 2024 (arXiv), 2025 (journal) | Yes | Hallucination, correctness plus groundedness | 17% / 17% / 33% / 43% |
| 3 | Thomson Reuters statement | 10 Jun 2024 | No | Internal accuracy | about 90% |
| 4 | Harvey BigLaw Bench | 29 Aug 2024 | No | Vendor benchmark | 74% of expert work product |
| 5 | Vals Legal AI Report | 27 Feb 2025 | Vendor-participating; LLM-as-judge | Task accuracy | 54% to 95% by task |
| 6 | Kalai et al. | 4 Sep 2025 | Yes (OpenAI authors) | Theory | none |
| 7 | Vals legal research benchmark | 23 Oct 2025 | Vendor-participating; blind human grading | Task accuracy | 79% to 81% vs lawyer 71% |
| 8 | Das et al., ClaimRAG-LAW | 20 May 2026 | Yes | Claim-level RAG | varying |
| 9 | Liu, Stammbach, Henderson | 19 Jun 2026 | Yes | Detection recall | GPT-5 84.4% recall, 55.0% F1 |
Appendix B: sources
- Dahl, Magesh, Suzgun and Ho, Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models: arXiv 2401.01301; Journal of Legal Analysis 16(1) 64 (2024)
- Magesh, Surani, Dahl, Suzgun, Manning and Ho, Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools: arXiv 2405.20362 (PDF read for this edition); Journal of Empirical Legal Studies (2025); Stanford HAI summary, AI on Trial (23 May 2024): hai.stanford.edu
- Thomson Reuters statement of 10 June 2024, as reported: Artificial Lawyer. Coverage of the study’s Westlaw results: LawSites, 4 June 2024
- Casetext marketing statement (2023), quoted verbatim at Magesh et al., arXiv 2405.20362, footnote 2
- Lou and Shin, Legal Reasoning Is Not Lawyering: Rethinking Legal Benchmarks for Pro Se Access to Justice: arXiv 2606.23716 (16 Jun 2026)
- LexisNexis press release, 25 October 2023: LexisNexis newsroom
- Harvey, Introducing BigLaw Bench (29 Aug 2024): harvey.ai
- Vals AI, Vals Legal AI Report (27 Feb 2025): vals.ai; VLAIR Legal Research (23 Oct 2025) as reported by LawSites
- Kalai, Nachum, Vempala and Zhang, Why Language Models Hallucinate: arXiv 2509.04664
- Das, Abualhaija and Bianculli, Fine-grained Claim-level RAG Benchmark for Law: arXiv 2605.21071
- Liu, Stammbach and Henderson, Who Checks the Citations? Benchmarking Legal Hallucination Detection: arXiv 2606.21155
- SafeLegalAI data: Incident Tracker · incidents.json · Legal AI Tech Tools · The sanctions ledger
Appendix C: changes to this report
- 5 September 2026: second edition. Expanded from two studies to a nine-row register; added the vendor claims table, the definitions section, the incident-record section and Appendix A. The July 2026 text’s figures for the two Stanford studies are unchanged; its statement that Ask Practical Law AI “hallucinated at a similar ~17% but gave incomplete answers on more than 60% of queries” is retained with the study’s accuracy figure (20%) added.
- 16 July 2026: first edition.
Reuse this research
Quote it, cite it, forward it — CC BY 4.0 for the data and figures; the linked official documents are the record. Suggested citation and a link to this exact report: