Could an AI's own draft of a paper review help expose reviews that a human handed off to an AI?
Can LLM-generated reference reviews detect machine-written peer review submissions?
This explores whether you can catch a reviewer who let an LLM write their peer review by generating an LLM's own review of the same paper and checking how closely the submitted review matches it.
This explores whether a 'reference review' written by an LLM can be used to spot reviews that were themselves machine-written. The short answer: this collection has no study that tests that exact method. It does have pieces that show why the idea might work, why it might not, and what conferences actually do instead.
The case for it working comes from an unexpected finding. AI reviewers show a 'hivemind' effect: across many papers, they agree with each other more than human reviewers agree with one another Can AI systems safely replace human peer reviewers?. That sameness is a weakness if you want AI to replace reviewers, but it's exactly what a reference-review detector would depend on. If machine reviews converge, a review that closely tracks what a model would say is suspicious. There's also evidence that models can learn to recognize their own text, and that this ability goes hand in hand with favoring it Do LLMs favor their own text because they recognize it?. So machine text seems to carry a recognizable signature.
The case against is practical. Humans can't reliably tell LLM-written research writing from human writing, and they tend to assume a person was involved either way Can readers tell LLM abstracts from human ones?. Rewording text is cheap: zero-shot rewrites shift AI scores without changing the substance Can AI systems safely replace human peer reviewers?, so a reviewer who lightly edits an LLM draft could easily drift away from any reference. Most real cases are also mixed rather than purely machine-written. In the ICML 2026 experiment, substantial shares of reviewers broke whichever LLM rule they were given, yet scores and decisions barely changed Does banning LLM use in peer review change review outcomes?. That makes 'was an LLM involved?' a blurrier and less consequential question than it sounds.
What conferences actually do is telling. ICLR 2026 treated detector flags as one weak signal passed to area chairs, not as an automatic verdict, because false positives were a real worry. The firm enforcement point turned out to be something else: fabricated references, meaning citations to papers that don't exist. These can be checked and confirmed, and papers with them were desk-rejected How can conferences detect and handle LLM misuse in peer review?. Sakana's AI-generated workshop paper was also caught out by a citation error after the fact Can AI-generated papers pass peer review undetected?. So the 'references' that reliably expose machine writing so far are bad citations, not comparison reviews.
The twist worth taking away: catching LLM use may matter less than catching bad reviews. Optional LLM feedback at ICLR 2025 led 27% of reviewers to make their reviews more specific Can LLM feedback help peer reviewers improve their own reviews?. The worry that LLM reviewers go easy on LLM-written papers disappears once you account for paper quality Do LLM reviewers actually favor LLM-written papers?. The field seems to be moving from 'who wrote this?' toward 'is this review any good?'
Sources 8 notes
AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.
Fine-tuning LLMs to recognize their own summaries increased their preference for those summaries in a linear relationship, suggesting recognition capability drives self-preference bias. The authors present this as initial causal evidence, not proof.
Readers with ML expertise struggle to identify LLM-generated content reliably, tending to assume human involvement across all abstract types. However, LLM-edited abstracts received highest clarity ratings and were preferred 55% of the time when authorship was disclosed.
A randomized experiment at ICML 2026 found that prohibiting LLM use versus allowing limited use barely changed paper scores, decisions, or reviewer confidence. Meanwhile, substantial fractions of reviewers broke whichever rule they were given.
Program chairs used imperfect detectors as one input for area chairs rather than automated filters, but desk-rejected papers with confirmed fabricated references as a tractable enforcement point. Multiple human review steps mitigated false positives.
Show all 8 sources
Sakana AI's end-to-end system produced a paper that scored 6.33 in double-blind ICLR 2025 workshop review, meeting acceptance thresholds, but was withdrawn under pre-agreed protocol. Authors later identified a citation error and judged none of three submissions suitable for main-track publication.
A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.
Across 125,000+ reviews, the apparent favoritism of LLM-assisted reviewers toward LLM papers disappears once paper quality is held constant. LLM papers cluster among weaker submissions, creating a spurious interaction driven by LLM reviewers' general leniency toward lower-quality work.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stop Automating Peer Review Without Rigorous Evaluation
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
- LLM-REVal: Can We Trust LLM Reviewers Yet?
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026
- LLM-Generated or Human-Written? Comparing Review and Non-Review Papers on ArXiv