A ban on AI in peer review barely moved the scores, and plenty of reviewers used AI anyway.
Do peer review policies banning LLM use actually change reviewer behavior and decisions?
This asks whether telling peer reviewers not to use LLMs actually changes what they do and how papers get scored, or whether the rule mostly exists on paper.
This asks whether a rule against LLM use in peer review changes what reviewers actually do and the scores papers end up with. The most direct evidence says it barely does. In a randomized experiment at ICML 2026, some reviewers were told not to use LLMs at all and others were allowed limited use. Paper scores, accept/reject decisions and reviewer confidence came out almost the same in both groups Does banning LLM use in peer review change review outcomes?. The reason matters: large shares of reviewers broke whichever rule they were given. So the experiment measured what a policy does, not what LLM use does. A ban doesn't neatly separate LLM users from non-users. It mostly changes who admits to it.
Enforcement shows the same gap. ICML planted hidden instructions in submitted PDFs as watermarks, which would show up in a review if someone pasted the paper into a model. This flagged 795 reviews, about 1%, and led to 497 desk rejections. The chairs themselves say this mostly catches careless use and misses anyone who stripped out or rewrote the planted text How many peer reviewers secretly used LLMs despite the ban?. Population-level estimates from earlier conferences put substantial LLM modification at 6.5% to 16.9% of reviews. It was more common among reviewers who were low-confidence, rushed or less engaged How much peer review text shows signs of LLM modification?. That means the ban mostly fails to reach the reviewers most likely to lean on a model. ICLR 2026 took a more practical line. Detector flags went to area chairs as one input among several, not as automatic verdicts. The hard enforcement was saved for something that can actually be checked: confirmed fabricated references How can conferences detect and handle LLM misuse in peer review?. Checking whether humans can spot AI text doesn't help much either. Readers with ML expertise couldn't reliably tell LLM abstracts from human ones Can readers tell LLM abstracts from human ones?.
The more interesting finding is that "did the reviewer use an LLM?" may be the wrong thing to police. What matters is how LLM-shaped judgment behaves. In simulations, LLM reviewers inflated scores for LLM-written papers and marked down human papers that made critical arguments Do LLM reviewers favor papers written by other LLMs?. Rewording a paper's rhetoric without changing its science moves AI review scores How much does rhetorical style shift AI review scores?. Simple rewrites raise AI scores by almost half a point, and AI reviewers agree with each other more than human reviewers do Can AI systems safely replace human peer reviewers?. Fake credentials and polished formatting reliably sway LLM judges Can LLM judges be fooled by fake credentials and formatting?. One twist: an analysis of more than 125,000 real reviews found that the apparent bias of LLM-assisted reviewers toward LLM papers disappears once paper quality is held constant. LLM-assisted reviewers are generally more lenient toward weaker papers, and LLM-written papers tend to be among the weaker ones Do LLM reviewers actually favor LLM-written papers?. So the real risk is less hidden favoritism and more a general softening of judgment.
That points to a different policy lever: managed use instead of prohibition. At ICLR 2025, a randomized trial offered reviewers optional feedback from Claude-based agents. 27% of reviewers revised their reviews, and blinded raters judged the revised versions more specific and informative Can LLM feedback help peer reviewers improve their own reviews?. Tightly structured uses, like a pipeline that breaks novelty assessment into claim extraction, related-work retrieval and comparison, reached 86% reasoning alignment with human reviewers Can structured pipelines make LLM novelty assessment reliable?. Put together, the evidence says bans mostly change what reviewers disclose rather than what they do. Measurable improvement came from building LLMs into review in narrow, checkable roles, not from trying to keep them out.
Sources 12 notes
A randomized experiment at ICML 2026 found that prohibiting LLM use versus allowing limited use barely changed paper scores, decisions, or reviewer confidence. Meanwhile, substantial fractions of reviewers broke whichever rule they were given.
Hidden-instruction watermarks planted in PDFs flagged about 1% of reviews under ICML's no-LLM rule, leading to 497 desk rejections. The chairs acknowledge the method catches mainly careless uses and misses reviewers who removed or rewrote the watermark.
Analysis of reviews from ICLR 2024, NeurIPS 2023, CoRL 2023, and EMNLP 2023 estimates this population share using distributional methods rather than per-review classification. Rates were higher in low-confidence, rushed, and less-engaged reviewers.
Program chairs used imperfect detectors as one input for area chairs rather than automated filters, but desk-rejected papers with confirmed fabricated references as a tractable enforcement point. Multiple human review steps mitigated false positives.
Readers with ML expertise struggle to identify LLM-generated content reliably, tending to assume human involvement across all abstract types. However, LLM-edited abstracts received highest clarity ratings and were preferred 55% of the time when authorship was disclosed.
Show all 12 sources
Simulated LLM reviewers gave higher scores to LLM-written papers and downrated human papers containing critical statements, while human annotators showed no such bias. The bias traces to preference for LLM writing style and aversion to critical framing.
Rewriting manuscripts' rhetoric while preserving scientific content moves LLM reviewer scores measurably. Evidence framing and novelty stance produce the largest contrasts; effects vary by reviewer model and interact with the paper's original quality tier.
AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.
Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.
Across 125,000+ reviews, the apparent favoritism of LLM-assisted reviewers toward LLM papers disappears once paper quality is held constant. LLM papers cluster among weaker submissions, creating a spurious interaction driven by LLM reviewers' general leniency toward lower-quality work.
A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.
A three-stage pipeline (extract claims, retrieve related work, compare) reached 86.5% reasoning alignment and 75.3% conclusion agreement with human reviewers on 182 ICLR submissions, outperforming holistic LLM baselines.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stop Automating Peer Review Without Rigorous Evaluation
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
- LLM-REVal: Can We Trust LLM Reviewers Yet?
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
- Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards
- Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026
- Pangram Predicts 21% of ICLR Reviews are AI-Generated