AI legal research tools sound safe, but does that just shift who gets caught faking citations, not how often it happens?
Do solo lawyers face different citation hallucination risks than large firms?
This explores whether a lawyer working alone is more exposed to AI-invented case citations than a big firm, or whether the two face different versions of the same risk.
This explores whether solo and small-firm lawyers are more exposed to AI-invented case citations than large firms, or whether each faces a different version of the risk. The headline number points one way. In one count of 114 US court cases with suspected AI errors, 90 percent involved solo or small firms and 56 percent involved plaintiff's counsel Do small law firms misuse AI more often than large ones?. That count only covers errors that got caught. It tells you who gets caught, not who misuses AI most often. Small practices usually have no second reader, no cite-checking staff, and often use general chatbots. Their mistakes are more likely to reach a judge unfiltered and in a form that is easy to spot.
Large firms face a quieter risk. They usually pay for legal research tools sold as safe, but a preregistered test found that Lexis+ AI, Westlaw AI-Assisted Research, and Ask Practical Law AI still produced hallucinations 17 to 33 percent of the time How often do legal AI tools actually hallucinate citations?. Because these are closed systems, outsiders can't check them independently. The danger moves from "I used ChatGPT and didn't check" to "I trusted the expensive tool my firm bought." That kind of error is easier to miss, because a plausible case from a trusted vendor gets less scrutiny than an obvious fake.
The errors are also hard to catch because of how people read citations. A study of 24,000 Search Arena interactions found that users favored answers with more citations almost as much when those citations were irrelevant as when they were relevant Do users trust citations more when there are simply more of them?. AI judges show the same weakness: fake references and polished formatting raise their scores whatever the content says Can LLM judges be fooled by fake credentials and formatting?. A brief full of citations looks authoritative to an overworked reader, whether that reader is a solo practitioner reviewing their own draft or a partner skimming an associate's work.
Some researchers argue that "hallucination" is the wrong word. The model produces true and false citations by exactly the same process, so the fix isn't better grounding but checking the output afterward Does calling LLM errors hallucinations point us toward the wrong fixes? Should we call LLM errors hallucinations or fabrications?. Seen that way, firm size matters mainly because it decides who can afford the checking. Courts don't make up the difference: when they find fabricated citations they rarely punish the fabrication itself, and sanctions depend on harm and intent Do courts actually sanction fabricated AI citations when detected?. Academic publishing offers a useful comparison. Hundreds of suspect citations turned up in accepted NeurIPS papers How many accepted conference papers contain hallucinated citations?. ICLR treated confirmed fake references as a clear-cut reason to reject a paper, and sent the fuzzier detector flags to human reviewers How can conferences detect and handle LLM misuse in peer review?. Courts could adopt a similar process, and it would protect solo lawyers better than relying on each lawyer's own discipline.
The corpus has no data on actual misuse rates by firm size, so the honest answer is that the risks differ in kind, and nobody has shown that solo lawyers misuse AI more often. Solo lawyers are more likely to have errors that go unchecked and get caught. Large firms are more likely to have errors that come from trusted vendor tools and slip through quietly.
Sources 9 notes
Of 114 US court cases with suspected AI errors, 90 percent involved solo or small firms and 56 percent involved plaintiff's counsel. However, this describes detected incidents, not base rates of misuse by firm size.
A preregistered evaluation found that Lexis+ AI, Westlaw AI-Assisted Research, and Ask Practical Law AI hallucinate between 17% and 33% of the time—far higher than vendors claim. Closed-system design prevents independent verification and accountability.
Analysis of 24,000 Search Arena interactions shows irrelevant citations boost user preference (β=0.273) nearly as much as relevant citations (β=0.285), indicating citation count functions as a decoupled trust heuristic.
Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.
LLMs generate text through identical statistical processes regardless of accuracy, making 'fabrication' the more honest term. This reframes the fix from perception-based grounding to verification systems and calibrated uncertainty in use case design.
Show all 9 sources
LLMs generate text through statistical token relationships without grounding in shared context. Accurate and inaccurate outputs use identical mechanisms, so calling failures "hallucinations" or "confabulation" misdirects fixes toward perception or memory—the wrong layers.
Five cases show courts found fabricated or suspected AI citations but imposed no dedicated penalties. Sanctions turned on discretion, demonstrated harm, and intent—not the hallucination itself.
GPTZero's citation checker flagged hundreds of potentially hallucinated citations across 4841 accepted NeurIPS 2025 papers. However, flagged citations require human verification to confirm hallucination, and the full verification rate across the full scan remains undisclosed.
Program chairs used imperfect detectors as one input for area chairs rather than automated filters, but desk-rejected papers with confirmed fabricated references as a tractable enforcement point. Multiple human review steps mitigated false positives.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools
- Who's Submitting AI-Tainted Filings in Court?
- A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models
- GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
- Ordinary, Reasonable Chatbots: Do AI Models Track Human Legal Judgments?
- Large Models of What? Mistaking Engineering Achievements for Human Linguistic Agency
- The Model Says Walk: How Surface Heuristics Override Implicit Constraints in LLM Reasoning
- The LLM Fallacy: Misattribution in AI-Assisted Cognitive Workflows