INQUIRING LINE

In a test at a major AI conference, many reviewers broke their assigned AI rules, and the rules barely changed scores.

How often do researchers violate rules about AI use in review?

This explores how often peer reviewers (and authors) break the rules conferences set about using AI, and what that rule-breaking tells us about whether such policies work at all.


This explores how often researchers break the rules on AI use in peer review, and what that says about the rules. The most direct evidence comes from a randomized experiment at ICML 2026. Reviewers were split into two groups: one was banned from using LLMs and the other was allowed limited use. Substantial fractions of reviewers broke whichever rule they were given Does banning LLM use in peer review change review outcomes?. The corpus summary doesn't give an exact violation rate, so if you want the numbers, go to that note. The surprising part is that the policy barely mattered. Paper scores, accept/reject decisions and reviewer confidence came out nearly the same under both rules.

The violations make more sense once you see how common AI use already is. A Frontiers survey of 1,645 researchers found that 53% of reviewers use AI tools, rising to 87% among early-career researchers. Most use it to draft reports or summarize papers How widely do peer reviewers actually use AI tools?. The same researchers said they want clearer policies. That points to a different reading of rule-breaking: some of it may come from confusion and habit rather than intent to cheat. If AI is already part of how half of reviewers work, a ban is asking many of them to change their routine.

Reviewers aren't the only ones bending the rules. Researchers found concealed instructions in 18 arXiv manuscripts telling AI reviewers to rate the paper favorably. These are text hidden from human readers but visible to a model Are hidden AI prompts in preprints a deceptive research practice?. The authors were betting that reviewers would quietly use AI anyway, so the authors' rule-breaking depends on the reviewers' rule-breaking. A survey of 230 publications describes this as a coupled arms race. Authors produce papers at scale, reviewers automate evaluation, authors manipulate the AI reviewers, and the cycle of defense and evasion continues Does AI create a coupled arms race in research production and review?.

The manipulation works because AI reviewers are easy to game. Simply rewriting a paper's text with AI raised AI-assigned scores by 0.45 points without improving the science. AI reviewers also tend to agree with each other more than human reviewers do Can AI systems safely replace human peer reviewers?. A related finding from writing research helps explain why hidden AI use spreads its flaws: people edited AI-written paragraphs only 23% of the time, and their edits left the text about 96% the same Do writers actually edit AI-generated text before publishing?. That study was about writers, not reviewers, but it suggests AI-drafted reviews may reach authors with little human filtering.

In short, rule-breaking is common enough that it showed up clearly in a controlled experiment. But the evidence suggests that the formal rule matters less than the incentives around it. The more useful question may be how review systems can hold up when AI use, permitted or not, is the norm. The corpus has only one study that measures violations directly, so a precise answer to "how often" is still thin.


Sources 6 notes

Does banning LLM use in peer review change review outcomes?

A randomized experiment at ICML 2026 found that prohibiting LLM use versus allowing limited use barely changed paper scores, decisions, or reviewer confidence. Meanwhile, substantial fractions of reviewers broke whichever rule they were given.

How widely do peer reviewers actually use AI tools?

Frontiers' May-June 2025 survey of 1,645 researchers found 53% of reviewers use AI tools, with adoption reaching 87% among early-career researchers. Most use AI for drafting reports or summarizing findings, and researchers express desire for clearer policies to guide more advanced applications.

Are hidden AI prompts in preprints a deceptive research practice?

Eighteen arXiv manuscripts contained concealed instructions directing AI reviewers to give positive assessments. The practice qualifies as questionable research conduct because concealment plus self-serving design violates ethics regardless of stated intent.

Does AI create a coupled arms race in research production and review?

A survey of 230 publications reveals production scaling, evaluation automation, manipulation, defenses, evasion, and ecosystem feedback as linked response relations among actors. Evidence is strongest for early stages and weakens toward long-horizon adaptation and feedback.

Can AI systems safely replace human peer reviewers?

AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.

Show all 6 sources
Do writers actually edit AI-generated text before publishing?

Writers edited AI-generated paragraphs only 23% of the time, with edits averaging 96% similarity to the original. This means AI's opinionated and distorted voice propagates with minimal human filtering before publication.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.