A top AI conference tested a full AI ban against limited use: scores barely moved, yet many reviewers broke whichever rule they got.
What happens when conferences enforce bans or limits on reviewer LLM use?
This explores what actually changes when AI conferences tell peer reviewers they can't use LLMs, or can use them only in limited ways: whether reviewers follow the rules, whether reviews and decisions change, and how organizers catch people who break them.
This explores what happens in practice when conferences ban or limit reviewers' use of LLMs: do reviewers comply, do outcomes shift, and how is the rule enforced? The surprising headline is that the choice of rule barely matters. ICML 2026 ran a randomized experiment that gave some reviewers a full ban and others permission for limited use. Paper scores, accept/reject decisions and reviewer confidence came out almost the same under both rules. Meanwhile, a substantial share of reviewers broke whichever rule they were given Does banning LLM use in peer review change review outcomes?. The policy debate tends to treat the wording of the rule as the lever, but this evidence suggests it isn't.
The rule-breaking makes more sense once you see where LLM use comes from. Before most of these policies existed, an estimated 6.5% to 16.9% of review text at major AI conferences had been substantially modified by an LLM. The rates were highest among reviewers who reported low confidence, wrote close to the deadline, or engaged less with the discussion How much peer review text shows signs of LLM modification?. That points to time pressure and shaky expertise as the driving forces, and a ban doesn't touch either one.
Enforcement is where the two big conferences have taken different routes. ICML hid instructions inside submission PDFs that act as watermarks: a reviewer who pastes the paper into an LLM may end up with telltale text in the review. This flagged 795 reviews, about 1% of the total, and led to 497 desk rejections. The chairs say openly that it mostly catches careless users and misses anyone who spotted and removed the hidden text or rewrote the output How many peer reviewers secretly used LLMs despite the ban?. ICLR 2026 treated detector flags as a reason to look closer, not as a verdict. Area chairs had to find concrete evidence before applying any sanction How can conferences detect and handle LLM misuse in peer review?. This trades wrongful penalties for extra hours of human review How do detection tools shape LLM use enforcement at ICLR?. The one thing ICLR did enforce automatically was fabricated references, because a citation that doesn't exist can be checked, while a sentence that merely sounds like an LLM cannot. Human judgment won't fill that gap either: even ML experts can't reliably tell LLM-written abstracts from human ones Can readers tell LLM abstracts from human ones?.
The main argument for restrictions has been fear of bias. In simulations, LLM reviewers gave higher scores to LLM-written papers and marked down human papers that took a critical stance Do LLM reviewers favor papers written by other LLMs?. Real conference data tells a more complicated story. Across more than 125,000 reviews, the apparent favoritism disappears once paper quality is held constant. LLM-assisted reviewers are generally more lenient toward weaker papers, and LLM-written papers tend to be among the weaker submissions Do LLM reviewers actually favor LLM-written papers?. The real problem may be leniency in rushed reviews, not machines favoring machines.
The unexpected part is that the most measurable gains came from giving reviewers more AI help, not less. At ICLR 2025, optional LLM feedback on draft reviews led 27% of reviewers to revise. Blinded raters judged the revised reviews more informative Can LLM feedback help peer reviewers improve their own reviews?. Narrow, structured tasks such as checking a paper's novelty against related work came close to human judgment Can structured pipelines make LLM novelty assessment reliable?. Taken together, the evidence suggests that the rules conferences write matter less than the tools they hand reviewers and the few problems that can be checked mechanically.
Sources 10 notes
A randomized experiment at ICML 2026 found that prohibiting LLM use versus allowing limited use barely changed paper scores, decisions, or reviewer confidence. Meanwhile, substantial fractions of reviewers broke whichever rule they were given.
Analysis of reviews from ICLR 2024, NeurIPS 2023, CoRL 2023, and EMNLP 2023 estimates this population share using distributional methods rather than per-review classification. Rates were higher in low-confidence, rushed, and less-engaged reviewers.
Hidden-instruction watermarks planted in PDFs flagged about 1% of reviews under ICML's no-LLM rule, leading to 497 desk rejections. The chairs acknowledge the method catches mainly careless uses and misses reviewers who removed or rewrote the watermark.
Program chairs used imperfect detectors as one input for area chairs rather than automated filters, but desk-rejected papers with confirmed fabricated references as a tractable enforcement point. Multiple human review steps mitigated false positives.
ICLR 2026 uses LLM detection tools only to triage papers for area chairs, who must find concrete evidence before sanctions are applied. This human gate protects against false positives by shifting costs from paper rejections to reviewer time.
Show all 10 sources
Readers with ML expertise struggle to identify LLM-generated content reliably, tending to assume human involvement across all abstract types. However, LLM-edited abstracts received highest clarity ratings and were preferred 55% of the time when authorship was disclosed.
Simulated LLM reviewers gave higher scores to LLM-written papers and downrated human papers containing critical statements, while human annotators showed no such bias. The bias traces to preference for LLM writing style and aversion to critical framing.
Across 125,000+ reviews, the apparent favoritism of LLM-assisted reviewers toward LLM papers disappears once paper quality is held constant. LLM papers cluster among weaker submissions, creating a spurious interaction driven by LLM reviewers' general leniency toward lower-quality work.
A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.
A three-stage pipeline (extract claims, retrieve related work, compare) reached 86.5% reasoning alignment and 75.3% conclusion agreement with human reviewers on 182 ICLR submissions, outperforming holistic LLM baselines.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stop Automating Peer Review Without Rigorous Evaluation
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
- LLM-REVal: Can We Trust LLM Reviewers Yet?
- Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards
- A Retrospective on the ICLR 2026 Review Process
- Pangram Predicts 21% of ICLR Reviews are AI-Generated