Rules on AI use in peer review often barely moved scores, partly because many reviewers simply broke them.
Why do peer review policies often fail to change actual review scores?
This explores why rules about how peer review should be done, such as banning or limiting reviewers' use of LLMs, often barely change the scores papers actually receive, and what that says about where review outcomes really come from.
This explores why changing the rules of peer review, especially rules about whether reviewers may use AI, so often leaves the actual scores unchanged. The clearest evidence comes from a randomized experiment at ICML 2026. Some reviewers were told not to use LLMs at all and others were allowed limited use. Paper scores, accept/reject decisions and reviewer confidence came out almost the same in both groups. The detail that explains much of this is that large numbers of reviewers broke whichever rule they were given Does banning LLM use in peer review change review outcomes?. A policy only on paper produces two groups that behave much more alike than their instructions suggest. The first reason, then, is simple: the policy was often never followed.
The second reason is less obvious. Scores may mostly reflect the paper being judged, and the review process may matter less. A study of more than 125,000 reviews found what looked like LLM-assisted reviewers favoring LLM-written papers. That apparent favoritism disappeared once paper quality was taken into account. LLM-written papers tended to be weaker, and LLM-assisted reviewers were generally more lenient toward weaker work. The two effects together created a pattern that looked meaningful but wasn't Do LLM reviewers actually favor LLM-written papers?. When paper quality drives most of the variation in scores, an effect that seems to come from the process can turn out to be a side effect of quality. A policy aimed at the process then has little left to change.
Third, the pressures that actually shape reviews are mostly things these policies don't touch. One model shows a feedback loop. Rising submission numbers overload unpaid reviewers, so journals recruit less qualified reviewers or pile more work on the existing ones. Accuracy drops, which encourages authors to submit more speculative papers, which raises submissions further Does peer review quality collapse under submission overload?. That model predicts falling accuracy, not a shift in average scores, and its real-world strength hasn't been measured yet. A position paper on AI conferences makes a related point: authors, reviewers and venues all share responsibility for review failures. It argues that fixes have to change incentives, for example by letting authors rate a review's quality before they see its verdict and by rewarding thorough reviewers Can two-stage review and badges fix AI conference peer review?. A rule about one tool used by one of those three parties leaves most of this in place. A survey of 230 publications describes the wider AI research-and-review ecosystem as a coupled arms race in which each defense prompts an evasion Does AI create a coupled arms race in research production and review?. The ICML rule-breaking fits that picture.
The surprise is that an intervention can succeed and still leave scores alone. At ICLR 2025, AI agents gave reviewers optional feedback on their draft reviews. Twenty-seven percent of reviewers revised their reviews, and blinded raters judged the revised reviews more informative and clearer Can LLM feedback help peer reviewers improve their own reviews?. The improvement showed up in the written review, not in the verdict. That suggests asking a different question: whether a policy makes reviews more useful to authors, rather than whether it moves scores. The review text appears much easier to influence than the final score. One caveat: none of these notes directly test why scores are hard to move. The explanations above are drawn from experiments, a model and position papers that each address part of the question.
Sources 6 notes
A randomized experiment at ICML 2026 found that prohibiting LLM use versus allowing limited use barely changed paper scores, decisions, or reviewer confidence. Meanwhile, substantial fractions of reviewers broke whichever rule they were given.
Across 125,000+ reviews, the apparent favoritism of LLM-assisted reviewers toward LLM papers disappears once paper quality is held constant. LLM papers cluster among weaker submissions, creating a spurious interaction driven by LLM reviewers' general leniency toward lower-quality work.
A two-journal model shows that rising submissions overtax unpaid reviewers, forcing journals to recruit less qualified reviewers or overload existing ones, which drops review accuracy and incentivizes authors to submit more speculatively, driving submissions higher. The mechanism is structural but its empirical strength remains to be measured.
Authors, reviewers, and venues all contribute to peer review failures at major AI conferences. A proposed two-stage system lets authors rate review quality before seeing verdicts, and a badge system rewards reviewer thoroughness, targeting measured biases like rating-length correlation.
A survey of 230 publications reveals production scaling, evaluation automation, manipulation, defenses, evasion, and ecosystem feedback as linked response relations among actors. Evidence is strongest for early stages and weakens toward long-horizon adaptation and feedback.
Show all 6 sources
A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stop Automating Peer Review Without Rigorous Evaluation
- Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
- Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026
- The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing
- LLM-REVal: Can We Trust LLM Reviewers Yet?