Skip to content
SafeLegalAI

72 on the record · +1 this week

reportAI Hallucinations

How often does legal AI hallucinate? What the studies actually measured

Stanford's research on legal-AI hallucination rates: general models fabricate at least 58% of the time; dedicated legal-research tools, 17% to 33%.

Daman Kaur

Two studies from Stanford’s RegLab give the best empirical answer to how often AI fabricates in legal work. General-purpose chatbots hallucinate on legal questions at least 58% of the time. Purpose-built legal-research tools do better, but still hallucinate between 17% and 33% of the time. Neither number supports using AI output without checking it. These figures are current as of the studies cited; both were peer-reviewed.

General-purpose models: at least 58%

In Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models (Journal of Legal Analysis, 2024), Dahl, Magesh, Suzgun, and Ho tested leading models on verifiable questions about federal cases. The result: when asked a direct, verifiable question about a randomly selected case, the models hallucinated between 58% and 88% of the time — GPT-4 at 58%, GPT-3.5 at 69%, PaLM 2 at 72%, and Llama 2 at 88%.

The study defined a hallucination as output not consistent with legal facts. The 58% figure is the best-case model; the paper’s own framing is that these systems hallucinate “at least 58%” of the time on such tasks.

The obvious rejoinder is that lawyers use dedicated tools with retrieval, not raw chatbots. The follow-up study — Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools (Journal of Empirical Legal Studies, 2025) — tested exactly those. It found the retrieval-augmented tools hallucinate less than general models, but not rarely:

  • Lexis+ AI answered about 65% of queries accurately and hallucinated roughly 17% of the time — more than one in six.
  • Westlaw AI-Assisted Research answered about 42% accurately and hallucinated roughly a third of the time — nearly twice the rate of the other tools tested.
  • Ask Practical Law AI hallucinated at a similar ~17% but gave incomplete answers on more than 60% of queries.

The study defined a hallucination along two axes: correctness (is the statement right?) and groundedness (does the cited source actually support it?). Thomson Reuters publicly disputed the query methodology; LexisNexis noted the product has been updated since testing. The direction of the finding is not seriously contested: retrieval reduces hallucination, it does not remove it.

What the numbers mean for practice

The evidence lines up with the incident record. Even the best legal-research tool tested hallucinated more than one in six times, and the general chatbots most litigants reach for hallucinated the majority of the time. That is why every court in the tracker, across twelve countries, converges on the same duty: verify each authority against the primary source before it goes before a court. The technology has improved; the obligation has not changed.

Sources