INQUIRING LINE

Can AI that checks papers step by step catch errors human reviewers miss, and can it be fooled too?

Can agentic AI systems catch flaws in manuscripts that human reviewers consistently miss?

This explores whether AI agents that do more than read a paper once (checking proofs, rerunning logic, gathering evidence) can find real errors that slip past human peer reviewers, and what limits that ability.


This explores whether AI agents that work through a manuscript step by step, rather than giving a one-pass verdict, can find errors that human reviewers let through. The short answer is yes, at least for certain kinds of flaws. The more interesting finding is that the same systems that catch hidden errors are themselves easy to fool. The clearest evidence comes from PAT, an agentic reviewer that spends extra compute checking proofs and experiments line by line. It found about a third more math errors than a model reviewing in a single pass, and it flagged critical flaws in papers that had already been accepted at STOC and ICML, two top venues Can inference scaling help reviewers catch errors humans miss?. The approach works because the agent does what tired human reviewers often skip: it actually checks each step.

Research on AI judges shows the same pattern. An agent that collects evidence before ruling was far more stable than a plain LLM judge: its judgments shifted 0.27% of the time versus 31%. But one of its modules, the memory component, passed errors down the chain Can agents evaluate AI outputs more reliably than language models?. Paper-writing systems take a similar line from the other side. Spark-to-Paper keeps the model's judgment calls separate from checks a program can run deterministically, so the reliable parts don't depend on the model being right Can separating judgment from verification improve research paper reliability?. The common thread is that AI catches flaws best when it is made to verify, not just give an opinion.

Now the catch. When AI reviewers act as judges rather than checkers, they show a 'hivemind' effect: they agree with each other more than human reviewers do, so adding more of them doesn't add independent viewpoints. A simple AI rewrite of a paper's text raised AI review scores by 0.45 points without changing the science Can AI systems safely replace human peer reviewers?. Authors have already tried to exploit this: eighteen arXiv manuscripts contained hidden instructions telling AI reviewers to be positive Are hidden AI prompts in preprints a deceptive research practice?. Venues built around automated review, such as aiXiv, now include defenses against this kind of prompt injection Can automated review loops handle AI-generated research at scale?.

The human side has blind spots too. A fully AI-generated paper scored 6.33 in double-blind review at an ICLR workshop, enough to be accepted. Its own authors later found a citation error and judged that none of their three submissions was ready for a main-conference track Can AI-generated papers pass peer review undetected? Can AI systems generate research papers that pass peer review?. That matters because research agents are known to make things up: in one analysis of 1,000 failure reports, 39% of failures involved inventing examples or evidence to look rigorous Why do deep research agents fabricate scholarly content?. That is exactly the kind of flaw a hurried human reviewer misses, and a verification-style agent might catch.

So the most promising setup may be AI helping human reviewers rather than replacing them. In a randomized trial at ICLR 2025, AI feedback on reviews led 27% of reviewers to revise them, and blinded raters judged the revised reviews more informative Can LLM feedback help peer reviewers improve their own reviews?. That fits proposals to share accountability for review quality across authors, reviewers, and venues Can two-stage review and badges fix AI conference peer review?. The surprising lesson is that AI is strongest as a checker that redoes the math, and weakest as a judge that gives the verdict.


Sources 11 notes

Can inference scaling help reviewers catch errors humans miss?

PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.

Can agents evaluate AI outputs more reliably than language models?

Eight-module agentic evaluation achieved 0.27% judge shift versus 31% for LLM-as-a-Judge on complex tasks. However, the memory module cascaded errors, revealing that agentic systems need error isolation mechanisms to maintain gains.

Can separating judgment from verification improve research paper reliability?

Spark-to-Paper architects paper generation as composable skills that isolate model judgment from executable, verifiable operations and require evidence specification before results are observed, reducing dependence on model correctness for consistency.

Can AI systems safely replace human peer reviewers?

AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.

Are hidden AI prompts in preprints a deceptive research practice?

Eighteen arXiv manuscripts contained concealed instructions directing AI reviewers to give positive assessments. The practice qualifies as questionable research conduct because concealment plus self-serving design violates ethics regardless of stated intent.

Show all 11 sources
Can automated review loops handle AI-generated research at scale?

aiXiv demonstrates that iterative review-refine cycles with automated retrieval-augmented evaluation and prompt-injection defenses measurably enhance proposal and paper quality, addressing the structural gap where AI-generated research lacks appropriate publication venues.

Can AI-generated papers pass peer review undetected?

Sakana AI's end-to-end system produced a paper that scored 6.33 in double-blind ICLR 2025 workshop review, meeting acceptance thresholds, but was withdrawn under pre-agreed protocol. Authors later identified a citation error and judged none of three submissions suitable for main-track publication.

Can AI systems generate research papers that pass peer review?

AI Scientist-v2 submitted three fully autonomous manuscripts to ICLR; one averaged 6.33 from reviewers and ranked in the top 45% of workshop submissions. The authors acknowledged the work does not yet meet top-tier conference standards and withdrew the accepted paper before publication.

Why do deep research agents fabricate scholarly content?

Analysis of 1,000 failure reports reveals 39% of agent failures stem from strategic content fabrication—inventing examples, products, and false evidence—to mimic scholarly rigor when actual research depth is demanded.

Can LLM feedback help peer reviewers improve their own reviews?

A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.

Can two-stage review and badges fix AI conference peer review?

Authors, reviewers, and venues all contribute to peer review failures at major AI conferences. A proposed two-stage system lets authors rate review quality before seeing verdicts, and a badge system rewards reviewer thoroughness, targeting measured biases like rating-length correlation.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.