Stricter AI rules for peer review barely moved paper scores in one randomized trial, and many reviewers broke the rule they got.
Do stricter AI policies actually change how reviewers score manuscripts?
This explores whether conference rules that ban or limit reviewers' use of LLMs actually change the scores and decisions papers receive, and what does move reviews if the rules don't.
This explores whether tightening the rules on AI use in peer review (a ban versus limited use) shows up in how papers get scored. The most direct evidence says it barely does. A randomized experiment at ICML 2026 gave reviewers different policies, either no LLM use or limited LLM use, and found almost no difference in paper scores, accept/reject decisions or reviewer confidence Does banning LLM use in peer review change review outcomes?. The second finding matters more: a large share of reviewers broke whichever rule they were given. So the policy may not change scores partly because it doesn't reliably change behavior.
That changes the question. If the rules don't move scores, what does? The corpus points to the manuscript itself and how it is written. When manuscripts were rewritten to change only their rhetoric, such as how evidence is framed and how strongly novelty is claimed, with the science left the same, LLM reviewer scores shifted measurably How much does rhetorical style shift AI review scores?. A separate study found that simple rewrites raised AI scores by about 0.45 points with no scientific improvement. It also found that AI reviewers agree with each other more than human reviewers do, a 'hivemind' effect Can AI systems safely replace human peer reviewers?. If reviewers quietly use AI despite a ban, these sensitivities follow them into the review. Some authors have already acted on this: eighteen arXiv manuscripts were found to contain hidden instructions telling AI reviewers to rate the paper positively Are hidden AI prompts in preprints a deceptive research practice?.
The approach with a measurable effect is the opposite of restriction. At ICLR 2025, a randomized trial offered reviewers optional feedback on their own reviews from Claude-based agents. Over a quarter of them (27%) revised their reviews, and blinded raters judged the revised versions more specific and informative Can LLM feedback help peer reviewers improve their own reviews?. Note that this changed the quality of the reviews, not necessarily the scores. Even so, using AI to improve reviewers' work changed what they did, while forbidding AI did not. A related line of work uses extra compute at review time to check proofs and experiments line by line. It caught serious errors in papers that had already passed human review at STOC and ICML Can inference scaling help reviewers catch errors humans miss?.
One way to make sense of this is to treat review as an arms race rather than a rule-following system. A survey of 230 publications describes six linked dynamics: more papers being produced, more automated evaluation, manipulation, defenses, evasion and feedback across the whole ecosystem. Each side's move prompts the other's response Does AI create a coupled arms race in research production and review?. Within that framing, a usage policy is one move that is easy to ignore. Proposals that change incentives try a different lever, for example letting authors rate review quality before they see the verdict, and giving badges for thorough reviewing Can two-stage review and badges fix AI conference peer review?.
The takeaway: stricter AI policies, as currently enforced, don't appear to change scores much. The reason isn't that AI has no effect on reviewing. Compliance is patchy, and the stronger influences are how a paper is written, what tools reviewers actually use and how reviewers are rewarded. The corpus has only one direct experiment on policy strictness, so this conclusion rests on a single venue and should be read as early evidence, not a settled result.
Sources 8 notes
A randomized experiment at ICML 2026 found that prohibiting LLM use versus allowing limited use barely changed paper scores, decisions, or reviewer confidence. Meanwhile, substantial fractions of reviewers broke whichever rule they were given.
Rewriting manuscripts' rhetoric while preserving scientific content moves LLM reviewer scores measurably. Evidence framing and novelty stance produce the largest contrasts; effects vary by reviewer model and interact with the paper's original quality tier.
AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.
Eighteen arXiv manuscripts contained concealed instructions directing AI reviewers to give positive assessments. The practice qualifies as questionable research conduct because concealment plus self-serving design violates ethics regardless of stated intent.
A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.
Show all 8 sources
PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.
A survey of 230 publications reveals production scaling, evaluation automation, manipulation, defenses, evasion, and ecosystem feedback as linked response relations among actors. Evidence is strongest for early stages and weakens toward long-horizon adaptation and feedback.
Authors, reviewers, and venues all contribute to peer review failures at major AI conferences. A proposed two-stage system lets authors rate review quality before seeing verdicts, and a badge system rewards reviewer thoroughness, targeting measured biases like rating-length correlation.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stop Automating Peer Review Without Rigorous Evaluation
- The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?
- Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards
- Pangram Predicts 21% of ICLR Reviews are AI-Generated
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review