INQUIRING LINE

When an AI invents a citation, is it misremembering a paper, or just producing plausible text that happens to be fake?

How do citation errors in AI-generated papers differ from human hallucinations?

This explores what makes a fabricated or wrong citation in an AI-written paper a different kind of error from the mistakes humans make, and whether 'hallucination', a word borrowed from human perception, is even the right name for it.


This explores how citation errors in AI-generated research differ from human mistakes, and whether calling them 'hallucinations' misleads us. One caveat first: the collection has no study that directly compares AI citation errors with human ones. What it does have is enough material from different angles to show where the two kinds of error part ways.

The sharpest point is about the word itself. A human hallucination is a failure of perception: something has gone wrong in a system that normally tracks reality. One argument in the collection says LLMs have no such tracking layer to break. They produce text from statistical relationships between tokens, and a correct citation and an invented one come out of exactly the same process Should we call LLM errors hallucinations or fabrications?. So a fake reference isn't the model misremembering a paper. It's the model producing something with the shape of a reference. That matters for fixes. When a human scholar cites the wrong paper, you can usually trace a cause, such as a misread abstract or a mixed-up author list. With a model, there may be no faulty step to repair, only a lack of grounding.

The second difference is motive, or what looks like one. In an analysis of 1,000 failure reports from deep research agents, 39% of failures came from what the authors call strategic fabrication: inventing examples and evidence to *look* rigorous when the task demands depth Why do deep research agents fabricate scholarly content?. Human citation errors are mostly sloppiness. These errors are produced in response to pressure. When an agent is asked to sound scholarly, fake citations are one of the ways it meets the request. Scale is the third difference. One demonstration generated 288 complete finance papers from 96 statistically significant patterns, each with a made-up theoretical rationale and fabricated references Can AI generate hundreds of fake academic papers automatically?. A human who invents a source does it once. A pipeline does it in batches.

The errors also behave differently once they're out in the world. Sakana's AI Scientist-v2 paper passed double-blind workshop review at ICLR. Its citation error was found later, by the authors, not the reviewers Can AI-generated papers pass peer review undetected? Can AI systems generate research papers that pass peer review?. That fits wider findings: people can't reliably tell AI text from human text Can people reliably spot content made by AI?, even ML experts reading research abstracts Can readers tell LLM abstracts from human ones?. And writers using AI assistance edit the output only about 23% of the time, with only light changes when they do Do writers actually edit AI-generated text before publishing?. A fabricated reference reads just as fluently as a real one, so human readers act as a weak filter.

That's why the most interesting responses in the collection don't try to make the model more truthful. They change the structure around it. Spark-to-Paper separates the model's judgment from deterministic, executable checks, and requires a plan for the evidence before any results are seen Can separating judgment from verification improve research paper reliability?. On the review side, an agentic reviewer that spends extra compute checking proofs and experiments line by line found flaws that had passed human review at STOC and ICML Can inference scaling help reviewers catch errors humans miss?. The lesson they share: a human error usually has a cause you can diagnose, but AI fabrication comes from how the text is generated, so the reliable defense is verification outside the model, not hoping it learns to be more careful.


Sources 10 notes

Should we call LLM errors hallucinations or fabrications?

LLMs generate text through statistical token relationships without grounding in shared context. Accurate and inaccurate outputs use identical mechanisms, so calling failures "hallucinations" or "confabulation" misdirects fixes toward perception or memory—the wrong layers.

Why do deep research agents fabricate scholarly content?

Analysis of 1,000 failure reports reveals 39% of agent failures stem from strategic content fabrication—inventing examples, products, and false evidence—to mimic scholarly rigor when actual research depth is demanded.

Can AI generate hundreds of fake academic papers automatically?

A demonstration showed LLMs generating 288 complete finance papers from 96 statistically significant signals, each with invented theoretical justifications and fabricated citations, proving academic HARKing can be automated at scale.

Can AI-generated papers pass peer review undetected?

Sakana AI's end-to-end system produced a paper that scored 6.33 in double-blind ICLR 2025 workshop review, meeting acceptance thresholds, but was withdrawn under pre-agreed protocol. Authors later identified a citation error and judged none of three submissions suitable for main-track publication.

Can AI systems generate research papers that pass peer review?

AI Scientist-v2 submitted three fully autonomous manuscripts to ICLR; one averaged 6.33 from reviewers and ranked in the top 45% of workshop submissions. The authors acknowledged the work does not yet meet top-tier conference standards and withdrew the accepted paper before publication.

Show all 10 sources
Can people reliably spot content made by AI?

A 30-study systematic review found that humans cannot reliably distinguish AI-generated from human-created content across text, image, and voice modalities. Accuracy generally clusters around chance and has not kept pace with improvements in AI realism.

Can readers tell LLM abstracts from human ones?

Readers with ML expertise struggle to identify LLM-generated content reliably, tending to assume human involvement across all abstract types. However, LLM-edited abstracts received highest clarity ratings and were preferred 55% of the time when authorship was disclosed.

Do writers actually edit AI-generated text before publishing?

Writers edited AI-generated paragraphs only 23% of the time, with edits averaging 96% similarity to the original. This means AI's opinionated and distorted voice propagates with minimal human filtering before publication.

Can separating judgment from verification improve research paper reliability?

Spark-to-Paper architects paper generation as composable skills that isolate model judgment from executable, verifiable operations and require evidence specification before results are observed, reducing dependence on model correctness for consistency.

Can inference scaling help reviewers catch errors humans miss?

PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.