SYNTHESIS NOTE
Topics›Domain Specialization›this note

How many peer reviewers secretly used LLMs despite the ban?

ICML used hidden watermarks in submission PDFs to detect LLM-written reviews submitted under a no-LLM policy. The question explores whether 795 flagged reviews represent the true scale of LLM use or only careless violations.

Synthesis note · 2026-10-06 · sourced from Domain Specialization

ICML's 2026 program chairs report that they caught LLM-written reviews by planting instructions in submission PDFs, and that the enforcement ended in 497 desk rejections. Of the reviews written under Policy A, the rule that no LLMs may be used, 795 (about 1% of all reviews) from 506 reviewers were flagged, and a human checked each flagged instance. The rejections follow from the rule's consequence: a flagged review by a reciprocal reviewer got that reviewer's own submission rejected, so the 497 papers correspond to 398 reciprocal reviewers. The chairs also removed 51 reviewers who had used LLMs in more than half of their reviews. The post was corrected on one point: the remaining 108 flagged reviewers were not reciprocal reviewers of active submissions.

The mechanism, which the chairs trace to recent work by Rao, Kumar, Lakkaraju, and Shah, turns the reviewer's own tool into the detector. The chairs built a dictionary of 170,000 phrases and sampled two per paper, a pair they say is less likely than one in ten billion. Each PDF carried instructions "visible only to an LLM" to include those two phrases in the review. A human reader would not see them, but a model reading the PDF would, and the chairs say the watermark "would subtly influence any review produced via an LLM." A review containing both phrases is therefore a fingerprint of a model that read the paper. Generic AI-text detectors were not used. The chairs report a family-wise false positive rate of 0.0001 and say every flag was inspected to rule out reviews that merely mentioned the watermark.

This is different evidence from the randomized experiment in Does banning LLM use in peer review change review outcomes?, which compared scores and decisions across policies and found noncompliance under both. The ICML post counts only what the watermark caught under the no-LLM rule, so it measures detected violations, not how often reviewers used LLMs. The same channel appears in Can LLM judges be fooled by fake credentials and formatting?, where hidden text steers a judge; here the hidden instruction is used as a detector instead. The excerpt also notes that the method "is not a difficult measure to circumvent," particularly once publicly known, which it was for almost the entire review period. Success rates were "over 80% for most models" in pre-deadline tests, but "not always."

The excerpt does not establish how common LLM use was among reviewers without the rule, or among those on Policy B, where limited use was allowed, so it cannot say whether 795 reviews is a large or small share of LLM-assisted reviewing. It also does not assess quality: "we are not making a judgment call about the quality of flagged reviews or the reviewers' intentions." The watermark misses anyone who removed it, rewrote the review, or used a model that ignored it. The implication, at the strength the evidence allows, is that 795 is a floor for what one method caught under one policy, not a prevalence estimate. Measuring noncompliance cleanly would take a design like the experiment's, not the watermark alone.

Inquiring lines that read this note 11

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Do restrictions on reviewer LLM use actually shape peer review behavior? How reliably can humans and AI detectors identify machine-generated text? How can we detect and account for LLM involvement in academic writing?

Related concepts in this collection 2

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 76 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

ICML says hidden-instruction watermarks flagged 795 reviews under its no-LLM rule and led to 497 desk rejections — possibly only the most careless uses