reportAI HallucinationsVendor Watch

How often does legal AI hallucinate? The measurement register, September 2026

Every published measurement of legal-AI error on one schema, set against the vendor claims: 17% to 33% for research tools, 43% to 88% for general models.

HarveyLexis+ AIOpenAI ChatGPT

Edited and verified by Cognesio LLP · updated

Researched with AI assistance · sources verified by Cognesio LLP · How this was made ↓

This report collects every published measurement of how often AI fabricates or misstates the law, in the United States and elsewhere, on one schema: who measured, what task, how many queries, what counted as a hallucination, which tool and version, when, and the rate. It then sets those measurements against what the vendors say about their own products. It is the second edition of a page first published on 16 July 2026, expanded from two studies to a register of nine entries and four vendor claims, and it will be updated as new measurements appear.

The short answer has not changed since 2024, because nobody has re-run the only rigorous test of the legal-research tools. On 202 preregistered legal queries in spring 2024, Lexis+ AI hallucinated in 17 percent of responses and answered 65 percent accurately; Westlaw’s AI-Assisted Research hallucinated in 33 percent and answered 42 percent accurately; GPT-4 hallucinated in 43 percent. Vendor-participating benchmarks since then report task accuracy of 54 to 95 percent, with different tasks and different graders, and the incident record shows 150 court decisions in which fabricated authority reached a court anyway, 93 of them dated 2026, to 3 September.

Key findings

  1. Nine entries, eight measurements and one theoretical account, have been published between January 2024 and August 2026, by four research groups, one benchmark company and two vendors. They use at least four different definitions of “hallucination” and are not comparable with each other.
  2. The only measurement of the major legal-research tools by researchers with no commercial interest (Magesh, Surani, Dahl, Suzgun, Manning and Ho, Journal of Empirical Legal Studies, 2025; queries run 22 March to 27 May 2024) found hallucination in 17 percent of Lexis+ AI responses, 17 percent of Ask Practical Law AI responses, 33 percent of Westlaw AI-Assisted Research responses and 43 percent of GPT-4 responses.
  3. Thomson Reuters answered that study on 10 June 2024 with an internal accuracy figure of “approximately 90%” for AI-Assisted Research, graded by two lawyers per question; the methods and the questions were not published, and the two numbers measure different things.
  4. LexisNexis’s launch claim of 25 October 2023, “hallucination-free linked legal citations”, is a claim about citation links, not about the correctness of the answer; the 2025 study found 17 percent of Lexis+ AI responses hallucinated under its definition, which includes correct citations that do not support the proposition.
  5. Two Vals AI benchmarks with vendor participation (27 February 2025 and 23 October 2025) report task accuracy of 54 to 95 percent for participating tools against lawyer baselines of 50 to 80 percent; Thomson Reuters and LexisNexis declined to enter the legal-research benchmark and vLex withdrew before the results were published, and the February grading was partly automated.
  6. General-purpose models measured on verifiable questions about US federal cases (Dahl, Magesh, Suzgun and Ho, Journal of Legal Analysis, 2024) hallucinated between 58 percent (GPT-4) and 88 percent (Llama 2) of the time.
  7. Two benchmarks and one position paper published in May and June 2026 address adjacent questions, not the rate itself: whether models can detect fabricated citations in a brief (GPT-5 reached 84.4 percent recall but 55.0 percent F1), whether retrieval-augmented legal systems can be evaluated at the claim level (the abstract motivates the benchmark on prior work finding that they still hallucinate at varying rates), and whether benchmark scores survive contact with a litigant’s own words (the authors argue current benchmarks cannot show that they do).
  8. The incident tracker’s 150 rows are a lagging indicator, not a rate: 92 of the 150 orders do not name the tool, and the rows dated 2026 (to 3 September) are 93 against 40 in 2025 because sweeps deepened, not because any measured rate rose.

Why this question

“How often does legal AI hallucinate” is the question behind every court rule on verification, every bar guidance note and every sanctions order in the incident tracker. It is also the question vendors answer with a marketing number and researchers answer with a benchmark number, and the two never describe the same thing. A general counsel deciding whether to allow a tool, a judge deciding whether to require certification, and a lawyer deciding how much to check all need the same three facts: what was measured, on what, and by whom with what interest.

The search record shows the demand: “ai with lowest hallucination rate” and “ai hallucination rate” carry 390 and 140 searches a month, and the pages that answer them are aggregators of vendor benchmarks. Nobody had put the measurements and the claims in one table with their methods exposed.

Method and data

A measurement is included if it reports a rate of hallucination, error or accuracy for a named AI system on a legal task, with a published method. Vendor blog posts count as claims, not measurements, unless the method is published. Each row of the register records: source and date; the systems tested; the task; the number of queries; who graded and how; the definition of hallucination or error; the reported rate. The register was compiled from the primary documents listed in Appendix B, each read for this edition (the ClaimRAG-LAW entry from its abstract only), and from the tools directory, which records each vendor’s published accuracy evidence with the field accuracyEvidence.

The incident data is the site’s incident collection, 150 records as of 5 September 2026, used only for the counts stated. Rates cannot be derived from it: the tracker records decisions that reached a court and were found, not the denominator of uses.

Disclosure: SafeLegalAI is published by Cognesio LLP, whose team also builds LegalAI Space, a governance product listed in the tools directory with an ownership label. LegalAI Space publishes no accuracy claim and is not in the register.

The register

#Measurement (date)SystemsTask and sizeGradingDefinitionReported rate
1Dahl, Magesh, Suzgun, Ho, Large Legal Fictions, arXiv 2 Jan 2024, J. Legal Analysis 2024GPT-4, GPT-3.5, PaLM 2, Llama 2Fourteen query types about a random sample of US federal cases (existence, citation, holding, authorship and others)Reference-based: answers checked against case metadata”Textual output that is not consistent with legal facts”Hallucination in 58% (GPT-4) to 88% (Llama 2) of responses on reference-based tasks; GPT-3.5 69%, PaLM 2 72%
2Magesh, Surani, Dahl, Suzgun, Manning, Ho, Hallucination-Free?, arXiv 30 May 2024, J. Empirical Legal Studies 2025Lexis+ AI, Westlaw AI-Assisted Research, Ask Practical Law AI, GPT-4202 preregistered open-ended legal research queries in four categories; run 22 Mar to 22 Apr 2024 (Lexis, Practical Law, GPT-4) and 23 to 27 May 2024 (Westlaw)Manual, two axes: correctness and groundedness; inter-rater validationIncorrect (misstates the law or a fact) or misgrounded (correct statement, cited source does not support it)Hallucinated: Lexis+ AI 17%, Ask Practical Law AI 17%, Westlaw AI-AR 33%, GPT-4 43%. Accurate (correct and grounded): Lexis+ AI 65%, Westlaw 42%, Practical Law 20%, GPT-4 49%. Incomplete: Practical Law 63%, Westlaw 25%, Lexis 18%, GPT-4 8%
3Thomson Reuters, statement by Mike Dahn, Head of Westlaw Product Management, 10 Jun 2024Westlaw Precision AI-Assisted Research”Hundreds of real-world legal research questions”; internalTwo lawyers per result, senior lawyer arbitrationNot published”Accuracy rate of approximately 90%” (internal; method and questions unpublished)
4Harvey, Introducing BigLaw Bench, 29 Aug 2024Harvey assistant models, foundation modelsReal-world legal work tasks (drafting, analysis); vendor-builtAnswer score (requirements met minus errors incl. hallucinations); source score (claims supported by an accurate source)Errors including hallucinated sources (document text or page number)Harvey outputs “74% of a final, expert lawyer-quality work product”; foundation models “consistently hallucinate sources” (vendor-run)
5Vals AI, Vals Legal AI Report, 27 Feb 2025CoCounsel (Thomson Reuters), Vincent AI (vLex), Harvey Assistant, Oliver (Vecflow); Lexis+ AI withdrew from the sections reportedSeven tasks (data extraction, document Q&A, summarisation, redlining, transcript analysis, chronology, EDGAR research); 240 questions across the seven tasks, over 500 graded samplesLLM-as-judge with lawyer baseline sourced by Cognia LawAccuracy per task, not hallucinationTask scores 53.6% to 94.8% (lowest: Vincent AI on redlining; highest: Harvey on document Q&A); lawyer baseline 50.3% to 80.2% by task; Harvey 94.8% on document Q&A; CoCounsel averaged 79.5% across four tasks; Vincent 53.6% to 72.7% and Oliver 55.2% to 74.0% across the tasks they entered
6Kalai, Nachum, Vempala, Zhang, Why Language Models Hallucinate, arXiv 4 Sep 2025GeneralTheoretical, with evaluation-incentive analysisNot applicableStatistical errors rewarded by evaluations that penalise abstentionNo rate; explains why rates persist: “training and evaluation procedures reward guessing over acknowledging uncertainty”
7Vals AI, VLAIR Legal Research, 23 Oct 2025Alexi, Counsel Stack, Midpage, ChatGPT (GPT-4)210 legal-research questions across nine research typesBlind grading by a consortium of law firms and academicsAccuracy of the answerLawyer baseline 71%; Alexi 80%, Counsel Stack 81%, Midpage 79%, ChatGPT 80%; Thomson Reuters, LexisNexis and vLex did not participate
8Das, Abualhaija, Bianculli, ClaimRAG-LAW, arXiv 20 May 2026 (rev. 14 Aug 2026)State-of-the-art legal and general RAG systems (unnamed in the abstract)Claim-level evaluation of retrieval and generation, French and EnglishClaim-level automated analysisClaim-level support of generated responsesNo rate in the abstract; the paper motivates itself on prior work showing that “RAG systems, whether general-purpose or legal-specific, still hallucinate at varying rates” and reports fine-grained retrieval and generation limitations (rates in the paper body, not read for this edition)
9Liu, Stammbach, Henderson, Who Checks the Citations?, arXiv 19 Jun 2026 (rev. 6 Aug 2026)Five models in agentic and non-agentic settings, plus Claude Code; best result from GPT-5Detecting injected fabricated citations in 1,300 brief excerptsAutomated against injected errorsDetection of a fabricated citation, not generation of oneGPT-5 84.4% recall, 55.0% F1 in the agentic setting, 15.3 steps per excerpt; “all models struggle with subtle error categories”

A tenth entry belongs in the register by omission: as of 5 September 2026 no independent re-test of Lexis+ AI, Westlaw AI-Assisted Research or CoCounsel on open-ended legal research had been published in the venues searched for this edition (arXiv, the Stanford RegLab publications page, Vals AI’s reports and the Journal of Empirical Legal Studies). Every number a buyer can cite for those products in September 2026 is either the 2024 study, a vendor’s internal figure, or a benchmark the vendor entered or declined.

Four definitions, four numbers

The register’s rates cannot be averaged because “hallucination” means four different things in it.

Not consistent with legal facts (row 1). The 2024 study asked models verifiable questions about real cases and scored the answer against the record. A wrong author, a wrong year or a non-existent case all count. The 58 to 88 percent range is a rate of false statements about cases the model was asked about directly.

Incorrect or misgrounded (row 2). The 2025 study added the second axis: a correct statement of law supported by a citation to a source that does not say it is a hallucination. This is the definition that matters for a brief, because a court checks the citation, not the proposition. It is also why the study’s 17 percent for Lexis+ AI is compatible with LexisNexis’s statement that its citations are linked and Shepardized: a real, valid, linked case can still be cited for a proposition it does not contain. The study gives Table 4 examples of Lexis+ AI citing the overruling case as authority for the overruled rule.

Task accuracy (rows 3, 5, 7). The vendor and benchmark figures score whether the answer to a task was right, as graded by a lawyer, two lawyers or a model. A response can be accurate on this definition and still contain a misgrounded citation; a response can be inaccurate without hallucinating anything. Thomson Reuters’ 90 percent and the Stanford 42 percent for the same product are not in conflict on their face, because one counts accurate answers to the vendor’s questions and the other counts hallucinated responses to preregistered questions the vendor did not choose.

Fabricated source (rows 4 and 9). BigLaw Bench’s source score and the June 2026 detection benchmark both treat the citation as the unit: does it exist, does it say that, can a model tell. This is the definition closest to what the incident tracker records, and it is the one on which the 2026 research says even GPT-5, working agentically over fifteen steps, misses a substantial share of subtle fabrications.

Vendor claims against the evidence

Vendor and productClaim (date, source)What the claim coversNearest measurementWhat the measurement found
LexisNexis, Lexis+ AI”hallucination-free linked legal citations” (press release, 25 Oct 2023); “minimizes the risk of invented content, or hallucinations, and checks all citations against Shepard’s”That linked citations resolve to real, validated documentsRow 2 (Mar to Apr 2024)17% of responses hallucinated (incorrect or misgrounded); 65% accurate; the highest accuracy of the tools tested. Lexis+ AI withdrew from the sections of the Feb 2025 Vals report
Thomson Reuters, Westlaw AI-Assisted Research”accuracy rate of approximately 90%” in internal testing; the product “can occasionally produce inaccuracies” (statement of 10 Jun 2024)Accuracy on the vendor’s own question set, two-lawyer gradedRow 2 (May 2024)33% of responses hallucinated; 42% accurate; “nearly twice as often as the other legal tools”
Thomson Reuters (Casetext), CoCounselCoCounsel “does not make up facts, or ‘hallucinate’” (Casetext, 2023, quoted in Magesh et al., n.2)Grounding in known data sourcesRow 5 (Feb 2025)Averaged 79.5% task accuracy across the four tasks entered in the Vals report; no independent hallucination measurement published
HarveyAssistant models produce “74% of a final, expert lawyer-quality work product”; foundation models “consistently hallucinate sources” (BigLaw Bench, 29 Aug 2024)Vendor-built benchmark of the vendor’s own productRow 5 (Feb 2025)Highest single task score in the Vals report (94.8%, document Q&A); tied the lawyer baseline on chronology; no independent hallucination rate published

Two patterns. First, every vendor claim is phrased about grounding or citation validity, and every independent measurement that found a problem found it in the proposition attached to a valid citation. Second, the two largest research vendors have each published one number since 2023, both internal, and neither has entered an independent benchmark of open-ended legal research: Thomson Reuters and LexisNexis did not enter the October 2025 Vals research benchmark, vLex withdrew from it, and Lexis+ AI withdrew from the reported sections of the February 2025 report. The tools directory records the published accuracy evidence per vendor: of 47 verified independent tools, 18 publish any accuracy evidence; 14 publish vendor evidence, 6 publish third-party evidence and 2 publish both (24 evidence entries in all, 15 vendor and 9 third-party).

The incident record as a lagging indicator

The tracker holds 150 decisions as of 5 September 2026 in which fabricated or misstated authority reached a court. Fifty-eight of the 150 orders name a tool. Counting rows that may name more than one: ChatGPT in 27, GPT-4o in 1, Claude in 5, Google products in 4, Westlaw or CoCounsel in 3 (one of them a row in which counsel denied using ChatGPT and named Westlaw’s tools), Microsoft Copilot in 2, a practice-management or specialist tool in 7 (LEAP, Visto.ai, Legal Genius, Strongsuit, ChatOn, a firm’s internal tool, Open Law), and 14 that say only that generative AI was used or suspected. Ninety-two do not identify the product.

Those counts are not rates. The denominator (how many filings used each tool) is unknown; the numerator is what courts wrote down and what the tracker found. The tool named most is the tool most people use. The three rows naming Westlaw or CoCounsel, one of them the Lacey v State Farm row naming CoCounsel, Westlaw Precision and Gemini together, are the only points where the register’s products and the incident record touch, and in each the court’s finding was that the lawyer did not check, not that the tool’s rate was known.

What the record does measure is that the mechanism the 2025 study described (a real citation attached to a proposition it does not support) is the one courts keep finding. The study’s typology of misgrounded responses reads as a catalogue of the tracker: holdings reversed, litigant arguments attributed to the court, overruled cases stated as good law without the citation. The rate at which that happens in practice remains unmeasured; the rate at which it is caught and sanctioned is in the sanctions ledger.

What a 2026 re-test would need

The 2025 study’s method supplies most of them, and its limitations section adds the hardest one: a preregistered query set, a definition that scores groundedness as well as correctness, manual grading with inter-rater checks, product versions and dates recorded, and access, because the vendors restrict their interfaces and the sample stayed at 202 queries. The June 2026 detection benchmark adds one: the graders themselves are now models, and the best of them misses fabrications that a human cite-checker would catch. A re-test that used a model to grade would inherit the error it was measuring.

Until one is published, the honest summary for a reader in September 2026 is the one this page carried in July: the best legal-research tool tested hallucinated in about one response in six on preregistered open-ended queries in spring 2024; the general chatbots most litigants use hallucinated in the majority of verifiable case questions in late 2023; the vendors report higher accuracy on their own tasks by their own graders; and no court in the tracker has treated any of those numbers as a reason not to check.

What to watch

The Stanford RegLab group has not announced a re-test; if one appears it will supersede row 2. Vals AI publishes new task reports through the year, and the participation list is the first thing to read in each. The detection benchmark (row 9) is the likeliest to be re-run against new model releases, and its recall figure is the number that bears on whether AI cite-checking can be trusted. On the incident side, the tracker’s tool-named share (58 of 150) will move as courts start requiring the tool to be identified under rules like Montana’s Fourth District Rule 3.G and Judge Graham’s standing order.

Three sentences journalists can quote

The only independent measurement of the major legal-research tools, on 202 preregistered queries in spring 2024, found that Lexis+ AI hallucinated in 17 percent of responses, Westlaw AI-Assisted Research in 33 percent and GPT-4 in 43 percent; it has not been repeated.

Thomson Reuters’ response was an internal accuracy figure of approximately 90 percent for the same product, graded by two lawyers on questions it has not published.

In the SafeLegalAI incident tracker, 92 of 150 court decisions on fabricated authority do not name the tool, so the record shows which products were admitted to, not which ones fail most.

Appendix A: the register as data

#SourceDateIndependent of vendorsRate typeHeadline figure
1Dahl et al., J. Legal AnalysisJan 2024 (arXiv), 2024 (journal)YesHallucination, reference-based58% to 88%
2Magesh et al., J. Empirical Legal StudiesMay 2024 (arXiv), 2025 (journal)YesHallucination, correctness plus groundedness17% / 17% / 33% / 43%
3Thomson Reuters statement10 Jun 2024NoInternal accuracyabout 90%
4Harvey BigLaw Bench29 Aug 2024NoVendor benchmark74% of expert work product
5Vals Legal AI Report27 Feb 2025Vendor-participating; LLM-as-judgeTask accuracy54% to 95% by task
6Kalai et al.4 Sep 2025Yes (OpenAI authors)Theorynone
7Vals legal research benchmark23 Oct 2025Vendor-participating; blind human gradingTask accuracy79% to 81% vs lawyer 71%
8Das et al., ClaimRAG-LAW20 May 2026YesClaim-level RAGvarying
9Liu, Stammbach, Henderson19 Jun 2026YesDetection recallGPT-5 84.4% recall, 55.0% F1

Appendix B: sources

Appendix C: changes to this report

  • 5 September 2026: second edition. Expanded from two studies to a nine-row register; added the vendor claims table, the definitions section, the incident-record section and Appendix A. The July 2026 text’s figures for the two Stanford studies are unchanged; its statement that Ask Practical Law AI “hallucinated at a similar ~17% but gave incomplete answers on more than 60% of queries” is retained with the study’s accuracy figure (20%) added.
  • 16 July 2026: first edition.

Reuse this research

Quote it, cite it, forward it — CC BY 4.0 for the data and figures; the linked official documents are the record. Suggested citation and a link to this exact report: