INQUIRING LINE

At ICML, banning reviewers from AI or allowing limited use produced nearly the same scores, and a substantial share broke their assigned rule.

How effective are journal policies restricting LLM use in peer review?

This explores whether rules limiting or banning LLM use by peer reviewers actually change anything: whether reviewers follow them, whether they can be enforced, and whether review outcomes differ. The corpus covers AI conferences such as ICML and ICLR rather than journals, so the answer comes from that setting.


This explores whether rules restricting reviewers' use of LLMs work, measured by compliance, enforcement and their effect on decisions. The corpus has no studies of journals, but AI conferences have tested these rules in practice, and the findings are surprising. The most direct evidence comes from a randomized experiment at ICML 2026. Some reviewers were told not to use LLMs at all, and others were allowed limited use. The two groups ended up with nearly the same paper scores, accept/reject decisions and reviewer confidence Does banning LLM use in peer review change review outcomes?. In addition, a substantial share of reviewers broke whichever rule they were assigned. So the policy did not change what reviewers did, and their behavior did not change the outcomes.

Enforcement is the weak link. ICML planted hidden instructions in submission PDFs, which act like a watermark: if a reviewer pastes the paper into a chatbot, a telltale phrase can appear in the review. This flagged 795 reviews, about 1% of the total, and led to 497 desk rejections. The chairs themselves say the method mostly catches careless users and misses anyone who removed the hidden instruction or rewrote the output How many peer reviewers secretly used LLMs despite the ban?. For comparison, an earlier analysis of ICLR, NeurIPS, CoRL and EMNLP reviews estimated that 6.5–16.9% of reviews showed substantial LLM modification. It also found the highest rates among low-confidence, last-minute and less-engaged reviewers How much peer review text shows signs of LLM modification?. The two numbers come from different venues and methods, so they aren't directly comparable. Still, the gap suggests that what gets caught is a small, visible fraction of actual use. It also points to the real problem: rushed reviewing, which an LLM ban alone doesn't fix.

ICLR 2026 took a more practical approach. Detector flags went to area chairs as one piece of evidence among several, not as automatic verdicts. Hard penalties were reserved for something that can actually be checked: fabricated references. Papers with confirmed hallucinated citations were desk-rejected How can conferences detect and handle LLM misuse in peer review?. This matters because detection is unreliable. Even ML experts can't consistently tell LLM-written text from human text Can readers tell LLM abstracts from human ones?. A rule against a fabricated reference can be enforced, while a rule against LLM-written text mostly can't.

The main concern behind these rules is bias, specifically that LLM reviewers favor LLM-written papers. Here the evidence is mixed. In simulations, LLM reviewers inflated scores for LLM-authored papers and marked down human papers that took a critical stance Do LLM reviewers favor papers written by other LLMs?. But across more than 125,000 real reviews, the apparent favoritism disappeared once paper quality was held constant. LLM-assisted reviewers were simply more lenient toward weaker papers in general, and LLM-written papers tended to be among the weaker ones Do LLM reviewers actually favor LLM-written papers?. Other risks are real, though. LLM judges can be fooled by fake authority signals and polished formatting Can LLM judges be fooled by fake credentials and formatting?, and polished prose no longer reliably signals a strong paper Does LLM writing assistance change how scientists publish?.

The less obvious conclusion is that the more promising policies shape how LLMs are used rather than trying to ban them. At ICLR 2025, an optional LLM tool gave reviewers feedback on their draft reviews. In a randomized trial, 27% of reviewers revised their reviews, and blinded raters judged the revisions more specific and informative Can LLM feedback help peer reviewers improve their own reviews?. A structured LLM pipeline for judging a paper's novelty agreed with human reviewers' reasoning 86% of the time Can structured pipelines make LLM novelty assessment reliable?. Evidence so far suggests bans are hard to enforce and change little. The questions that remain open are where in the review process LLMs belong and which failures, like fabricated references, are worth policing.


Sources 11 notes

Does banning LLM use in peer review change review outcomes?

A randomized experiment at ICML 2026 found that prohibiting LLM use versus allowing limited use barely changed paper scores, decisions, or reviewer confidence. Meanwhile, substantial fractions of reviewers broke whichever rule they were given.

How many peer reviewers secretly used LLMs despite the ban?

Hidden-instruction watermarks planted in PDFs flagged about 1% of reviews under ICML's no-LLM rule, leading to 497 desk rejections. The chairs acknowledge the method catches mainly careless uses and misses reviewers who removed or rewrote the watermark.

How much peer review text shows signs of LLM modification?

Analysis of reviews from ICLR 2024, NeurIPS 2023, CoRL 2023, and EMNLP 2023 estimates this population share using distributional methods rather than per-review classification. Rates were higher in low-confidence, rushed, and less-engaged reviewers.

How can conferences detect and handle LLM misuse in peer review?

Program chairs used imperfect detectors as one input for area chairs rather than automated filters, but desk-rejected papers with confirmed fabricated references as a tractable enforcement point. Multiple human review steps mitigated false positives.

Can readers tell LLM abstracts from human ones?

Readers with ML expertise struggle to identify LLM-generated content reliably, tending to assume human involvement across all abstract types. However, LLM-edited abstracts received highest clarity ratings and were preferred 55% of the time when authorship was disclosed.

Show all 11 sources
Do LLM reviewers favor papers written by other LLMs?

Simulated LLM reviewers gave higher scores to LLM-written papers and downrated human papers containing critical statements, while human annotators showed no such bias. The bias traces to preference for LLM writing style and aversion to critical framing.

Do LLM reviewers actually favor LLM-written papers?

Across 125,000+ reviews, the apparent favoritism of LLM-assisted reviewers toward LLM papers disappears once paper quality is held constant. LLM papers cluster among weaker submissions, creating a spurious interaction driven by LLM reviewers' general leniency toward lower-quality work.

Can LLM judges be fooled by fake credentials and formatting?

Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.

Does LLM writing assistance change how scientists publish?

Across 2.1M preprints, scientists using LLMs showed 23.7–89.3% higher publication rates. Simultaneously, the correlation between complex prose and paper quality reversed, suggesting polish no longer reliably indicates scientific merit.

Can LLM feedback help peer reviewers improve their own reviews?

A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.

Can structured pipelines make LLM novelty assessment reliable?

A three-stage pipeline (extract claims, retrieve related work, compare) reached 86.5% reasoning alignment and 75.3% conclusion agreement with human reviewers on 182 ICLR submissions, outperforming holistic LLM baselines.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.