INQUIRING LINE

Court records of caught AI errors can't say whether ChatGPT turns up more than legal tools, just who gets caught.

Does ChatGPT appear in more cases than other legal AI tools?

This explores whether ChatGPT shows up more often than purpose-built legal AI tools in court cases where AI-generated errors were caught. The corpus can't settle the tool-by-tool count, but it does show who gets caught and why general chatbots are easy to over-trust.


This explores whether ChatGPT shows up more often than specialized legal AI products in court cases involving AI mistakes. The short answer is that the collection doesn't have a tool-by-tool count. Its closest evidence is a study of 114 US court cases with suspected AI errors, and the finding it reports is about who files them, not which product produced them. Ninety percent involved solo or small firms, and 56 percent involved plaintiff's counsel. The author warns that these are detected incidents, not true rates of misuse Do small law firms misuse AI more often than large ones?. That warning applies to your question too. A tally of tools named in caught cases would partly measure which tools get caught and admitted to, not which ones err most.

The firm-size pattern still hints at an answer. Small practices are the least likely to pay for enterprise legal research platforms, so a general chatbot is the obvious tool for them. If the cases cluster where general chatbots are the default, ChatGPT's prominence may say more about who can afford what than about ChatGPT being uniquely unreliable at law.

The corpus also explains why a general chatbot is easy to over-trust. A focus-group study found that people trust ChatGPT because of its conversational back-and-forth, its speed and its tidy formatting, not because they've checked its accuracy Does conversational style actually make AI more trustworthy?. Research on AI graders shows the same weakness from the other side: LLM judges give higher scores to answers with fake references and polished formatting, whatever the content says Can LLM judges be tricked without accessing their internals?. A confident brief full of authoritative-looking citations is exactly the kind of output that gets past a busy reader, whether that reader is a model or a lawyer.

Two more findings complicate the idea that the tool is the whole problem. In a randomized trial, GPT-4 made law students much faster but improved the quality of their analysis only slightly, mostly for weaker students Does GPT-4 actually improve the quality of legal analysis?. The speed is obvious while the quality gain is small, and that mismatch invites people to skip checking the work. Separately, a benchmark of Supreme Court overruling cases found that models do worse on older precedent because their training data over-represents recent cases Why do language models struggle with historical legal cases?. Older case law is exactly where a model is most likely to make up a plausible citation.

The takeaway you might not expect is that 'which tool appears most' is probably the wrong question. Tool counts in court records mostly mirror market share and who gets caught. What the corpus does show is how failures happen: fluent chat that earns trust it hasn't earned, big speed gains alongside small quality gains, and weak coverage of older law. Those apply to any LLM a lawyer uses, branded legal product or not. The collection doesn't directly compare ChatGPT with tools like Westlaw's or Lexis's AI features, so that comparison is still open.


Sources 5 notes

Do small law firms misuse AI more often than large ones?

Of 114 US court cases with suspected AI errors, 90 percent involved solo or small firms and 56 percent involved plaintiff's counsel. However, this describes detected incidents, not base rates of misuse by firm size.

Does conversational style actually make AI more trustworthy?

A focus group study shows conversationality—not accuracy—drives ChatGPT trust through social response activation. Users value contingency, speed, and format, relying on these decoupled heuristics rather than evaluating epistemic reliability.

Can LLM judges be tricked without accessing their internals?

Research shows LLM evaluators systematically score higher when responses include fake references or rich formatting, independent of content quality. These biases are exploitable without model access, undermining AI benchmark credibility.

Does GPT-4 actually improve the quality of legal analysis?

A randomized controlled trial found GPT-4 saved consistent time across all skill levels and improved quality unevenly, with largest gains for weaker students. Speed effects were large and consistent; quality effects were small and concentrated among lower-skilled participants.

Why do language models struggle with historical legal cases?

Supreme Court overruling benchmark (236 pairs) reveals era sensitivity: models perform worse on historical cases than modern ones. Root cause is training corpus over-representation of recent cases, creating shallower representations of older precedent.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.