INQUIRING LINE

Nobody has measured how often AI-written papers slip factual errors past reviewers, but the few documented cases hint at who actually catches them.

How often do AI systems produce papers with undetected factual errors?

This explores how often AI-written research papers contain factual mistakes that slip past reviewers. The corpus has no overall rate, but it shows why errors get through and where they have actually been caught.


This explores how often AI-generated research papers contain factual errors that nobody catches. The short answer is that the collection has no rate to give you. Nobody here has measured it across many papers. What the corpus offers instead is a handful of telling cases and a clearer picture of why the number is hard to get.

The best-documented case is small. Sakana AI's AI Scientist-v2 sent three fully machine-written papers to an ICLR 2025 workshop. One scored 6.33 in double-blind review, enough to be accepted, and was withdrawn under a protocol agreed in advance Can AI-generated papers pass peer review undetected?. The detail worth noticing is who found the mistake. The reviewers didn't flag the citation error; the authors found it afterward. They also judged that none of the three papers was good enough for the main conference Can AI systems generate research papers that pass peer review?. So in the one real test of AI papers under peer review, the paper that passed carried an error the reviewers missed. That's a sample of one, but it points the wrong way.

A darker demonstration shows what happens at volume. An LLM pipeline produced 288 complete finance papers from 96 statistically significant signals. Each paper came with an invented theory to explain its result and fabricated citations Can AI generate hundreds of fake academic papers automatically?. Here the errors aren't accidents. They are what the method produces: the theory is written after the result is known, and the citations are made up. If AI models are then used to screen papers like these, there is a further weakness. LLM judges give higher scores to responses that include fake references or polished formatting, whatever the content Can LLM judges be tricked without accessing their internals?. The signs of rigor are exactly what fools the checker.

The surprising lateral finding is that human review was already missing errors in human-written work. PAT, an agentic reviewer that spends extra compute checking proofs and experiments line by line, found critical flaws in papers at STOC and ICML that had already passed expert review Can inference scaling help reviewers catch errors humans miss?. So 'undetected' says as much about the checking as about the AI. Agent judges that collect evidence as they go are far more stable than plain LLM judges, but one study found errors in the memory module spreading through the rest of the system Can agents evaluate AI outputs more reliably than language models?. Other work tries to design the problem out. Spark-to-Paper keeps the model's judgment calls separate from checks that run deterministically. It also requires authors to state what evidence would count before they see their results Can separating judgment from verification improve research paper reliability?.

The human side of the pipeline offers little protection either. Writers edited AI-drafted paragraphs only 23% of the time, and their edits left the text 96% the same Do writers actually edit AI-generated text before publishing?. Showing readers the AI's reasoning tends to make them accept its answer whether or not it is right. Only explanations that argue both for and against the answer actually help people spot mistakes Do explanations actually help users spot AI mistakes?. That brings back the missing number. Tools exist for measuring parts of the problem, such as whether errors are visible, contained, or reversible. None of them yet measures how errors pass through a whole system of authors, reviewers and institutions How can we measure whether AI errors stay visible and recoverable?. The real gap isn't that AI papers have errors. It's that we lack the means to count how many get through.


Sources 10 notes

Can AI-generated papers pass peer review undetected?

Sakana AI's end-to-end system produced a paper that scored 6.33 in double-blind ICLR 2025 workshop review, meeting acceptance thresholds, but was withdrawn under pre-agreed protocol. Authors later identified a citation error and judged none of three submissions suitable for main-track publication.

Can AI systems generate research papers that pass peer review?

AI Scientist-v2 submitted three fully autonomous manuscripts to ICLR; one averaged 6.33 from reviewers and ranked in the top 45% of workshop submissions. The authors acknowledged the work does not yet meet top-tier conference standards and withdrew the accepted paper before publication.

Can AI generate hundreds of fake academic papers automatically?

A demonstration showed LLMs generating 288 complete finance papers from 96 statistically significant signals, each with invented theoretical justifications and fabricated citations, proving academic HARKing can be automated at scale.

Can LLM judges be tricked without accessing their internals?

Research shows LLM evaluators systematically score higher when responses include fake references or rich formatting, independent of content quality. These biases are exploitable without model access, undermining AI benchmark credibility.

Can inference scaling help reviewers catch errors humans miss?

PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.

Show all 10 sources
Can agents evaluate AI outputs more reliably than language models?

Eight-module agentic evaluation achieved 0.27% judge shift versus 31% for LLM-as-a-Judge on complex tasks. However, the memory module cascaded errors, revealing that agentic systems need error isolation mechanisms to maintain gains.

Can separating judgment from verification improve research paper reliability?

Spark-to-Paper architects paper generation as composable skills that isolate model judgment from executable, verifiable operations and require evidence specification before results are observed, reducing dependence on model correctness for consistency.

Do writers actually edit AI-generated text before publishing?

Writers edited AI-generated paragraphs only 23% of the time, with edits averaging 96% similarity to the original. This means AI's opinionated and distorted voice propagates with minimal human filtering before publication.

Do explanations actually help users spot AI mistakes?

Reasoning traces and post-hoc explanations increase user acceptance of AI answers regardless of correctness, engendering false trust. Only dual explanations presenting arguments for and against the answer genuinely help users distinguish correct from incorrect outputs.

How can we measure whether AI errors stay visible and recoverable?

Partial instruments exist for individual conditions in isolated settings, but none measures the full socio-technical system the paper identifies as necessary. Visibility has a model-side measure (chain-of-thought disclosure), containment has incident-level counts, and recoverability has rollback timing, yet none bridges all four or captures human-institution factors.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.