When a court catches an AI making up fake case citations, how is that different from a lawyer's honest typo?
How do courts distinguish between AI hallucinations and ordinary typographical errors?
This explores how judges tell a fabricated AI citation apart from an honest clerical mistake, and whether that difference changes how they respond.
This explores how judges tell a fabricated AI citation apart from an honest clerical mistake, and whether the difference matters to the outcome. The collection has no paper that describes a formal legal test for separating the two. What it does show is more surprising: the courts that have been studied mostly don't treat the line between them as the deciding question. In five cases where courts found fabricated or suspected AI citations, none imposed a penalty for the hallucination itself. Sanctions depended on the judge's discretion, on whether the error caused real harm, and on whether the lawyer acted in bad faith Do courts actually sanction fabricated AI citations when detected?. So in practice, courts handle AI errors the way they already handle careless filings. They ask about intent and consequences, not about which tool produced the mistake.
There's a good reason the 'how did this error happen' question is hard to answer from the outside. One line of work argues that 'hallucination' is the wrong word. A language model produces accurate and inaccurate text through exactly the same statistical process, so its errors are better called fabrication than a misperception or a slip of memory Should we call LLM errors hallucinations or fabrications?. That marks a real difference from a typo. A typo is a slip in recording something real: the case exists, but the page number is wrong. A fabricated citation can point to a case that never existed, written in the same confident, well-formatted style as a real one. Formal proofs suggest no computable model can avoid this completely, which is why outside checks are necessary rather than optional Can any computable LLM truly avoid hallucinating?.
The legal industry's own tools don't remove the problem. A preregistered evaluation found that Lexis+ AI, Westlaw's AI-assisted research and Ask Practical Law AI produced hallucinations in 17% to 33% of answers, despite being marketed as 'hallucination-free'. Because these systems are closed, outsiders can't easily verify where an error came from How often do legal AI tools actually hallucinate citations?. A lawyer who files a bad citation from one of these tools may honestly not know whether they introduced the error or the software did.
The broader research explains why judges catch these errors by checking the citation, not by reading closely. People perform at roughly chance when trying to spot AI-generated content Can people reliably spot content made by AI?. AI text differs from human writing in ways statistics can measure, yet even trained linguists can't perceive those differences Can humans detect AI text if machines can measure it?. In an 81-person study, readers fell for fluent fabrications as readily as for true statements, until an interface showed them which claims were backed by sources Can readers tell truth from fabrication without evidence signals?. Fake citations also work on AI evaluators. LLM 'judges' give higher scores to answers that include fake references, which shows how much citations alone can make a weak argument look credible Can LLM judges be fooled by fake credentials and formatting?.
The takeaway: the useful line isn't 'AI error versus typo' but 'mistake in recording a real source versus a source that doesn't exist'. The only reliable way to tell them apart is to look the source up. Courts seem to have accepted this, focusing on whether lawyers checked their work and whether anyone was harmed. If you're interested in where this goes next, the provenance study suggests the real fix is interfaces that show which claims are backed by sources, not better detectors.
Sources 8 notes
Five cases show courts found fabricated or suspected AI citations but imposed no dedicated penalties. Sanctions turned on discretion, demonstrated harm, and intent—not the hallucination itself.
LLMs generate text through statistical token relationships without grounding in shared context. Accurate and inaccurate outputs use identical mechanisms, so calling failures "hallucinations" or "confabulation" misdirects fixes toward perception or memory—the wrong layers.
Three formal theorems prove that any computable LLM must hallucinate on infinitely many inputs, and internal mechanisms like self-correction cannot eliminate this mathematical constraint. External safeguards are therefore necessary, not optional.
A preregistered evaluation found that Lexis+ AI, Westlaw AI-Assisted Research, and Ask Practical Law AI hallucinate between 17% and 33% of the time—far higher than vendors claim. Closed-system design prevents independent verification and accountability.
A 30-study systematic review found that humans cannot reliably distinguish AI-generated from human-created content across text, image, and voice modalities. Accuracy generally clusters around chance and has not kept pace with improvements in AI realism.
Show all 8 sources
LLM-generated text differs significantly on six lexical diversity dimensions, confirmed through statistical analysis across multiple models. Yet human judges, including trained linguists, cannot reliably detect these differences—and newer models diverge further while becoming harder to spot.
In an 81-person study, participants given no provenance cues showed no significant truth discernment (p = .43), falling for fluent hallucinations as readily as ground truth. An idealized Provenance Density interface showing verified claims restored a +4.15 point gap (p < .001).
Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Is it Cake or is it AI? A Systematic Review of Human Uncertainty in Distinguishing Generative Artificial Intelligence Content
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
- A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models
- Hallucination is Inevitable: An Innate Limitation of Large Language Models
- A comprehensive taxonomy of hallucinations in Large Language Models
- Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools
- Measuring AI "Slop" in Text
- Do LLMs produce texts with "human-like" lexical diversity?