INQUIRING LINE

A top AI conference hid secret instructions in submitted papers to catch AI-written reviews, but no one knows how many slipped past.

Did reviewers successfully circumvent ICML's hidden-instruction watermark detection method?

This explores whether peer reviewers at ICML got around the conference's trick of hiding instructions inside submitted PDFs to catch reviews written by an LLM, and what the corpus can and can't say about how many slipped through.


This explores whether reviewers at ICML got past the conference's method for catching LLM-written reviews: instructions hidden inside the submitted PDFs that an LLM would follow but a human reader would never see. The short answer is that nobody can count the ones who got past it. The method caught 795 reviews, about 1% of the total, and led to 497 desk rejections. The program chairs say plainly that it mostly catches careless use. A reviewer who spotted the hidden instruction and deleted it, or who rewrote the LLM's output by hand, would leave no trace How many peer reviewers secretly used LLMs despite the ban?. So some reviewers almost certainly evaded it, but no one knows how many. The 795 figure is a minimum. It isn't an estimate of how many reviewers actually used LLMs.

The less obvious point is that the same trick runs in both directions. ICML used hidden instructions to catch reviewers. Some authors used the same mechanism to manipulate them: eighteen arXiv manuscripts were found carrying concealed prompts telling AI reviewers to give positive assessments Are hidden AI prompts in preprints a deceptive research practice?. Hidden text in a PDF has become a contested space. Conference organizers use it to police reviewers, and authors use it to sway whatever AI is reading the paper. A reviewer pasting a PDF into an LLM is exposed to both at once, and the self-serving version is now treated as a breach of publication ethics.

If the corpus can't measure evasion directly, can other detection methods fill the gap? Only weakly. One study claims that heavily rewritten text also evades AI-text detectors, but it never actually ran detectors. It only showed that rewrites erase stylistic fingerprints Do rewrites that hide authorship also fool AI detectors?. That is the same move a careful reviewer would make to beat ICML's watermark, so a second layer of detection probably wouldn't catch them either, though this is untested. ICLR 2026 took a different approach. It treated detector flags as one signal for human area chairs rather than an automatic verdict, and it reserved hard enforcement for something easy to verify: fabricated references How can conferences detect and handle LLM misuse in peer review?. That approach accepts that clever evaders will get through and goes after the errors LLMs leave behind instead.

There is a further reason the evasion question matters. If reviewers quietly hand judgment to an LLM, they inherit its weaknesses as a judge. LLM evaluators reliably score work higher when it includes fake references or polished formatting, whatever the actual quality Can LLM judges be fooled by fake credentials and formatting? Can LLM judges be tricked without accessing their internals?. So a review that escaped ICML's net is not just a broken rule. It may be a review that an author can steer with cosmetic tricks.

On the gap itself: the corpus has the chairs' admission that evasion was possible and likely. It has no study measuring how many ICML reviewers removed or rewrote the watermark, and by design that number can't be observed.


Sources 6 notes

How many peer reviewers secretly used LLMs despite the ban?

Hidden-instruction watermarks planted in PDFs flagged about 1% of reviews under ICML's no-LLM rule, leading to 497 desk rejections. The chairs acknowledge the method catches mainly careless uses and misses reviewers who removed or rewrote the watermark.

Are hidden AI prompts in preprints a deceptive research practice?

Eighteen arXiv manuscripts contained concealed instructions directing AI reviewers to give positive assessments. The practice qualifies as questionable research conduct because concealment plus self-serving design violates ethics regardless of stated intent.

Do rewrites that hide authorship also fool AI detectors?

The paper asserts that rewritten messages evade AI-text detectors but provides no detector experiments, only attribution results showing stylistic convergence. The double erasure claim needs direct empirical testing.

How can conferences detect and handle LLM misuse in peer review?

Program chairs used imperfect detectors as one input for area chairs rather than automated filters, but desk-rejected papers with confirmed fabricated references as a tractable enforcement point. Multiple human review steps mitigated false positives.

Can LLM judges be fooled by fake credentials and formatting?

Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.

Show all 6 sources
Can LLM judges be tricked without accessing their internals?

Research shows LLM evaluators systematically score higher when responses include fake references or rich formatting, independent of content quality. These biases are exploitable without model access, undermining AI benchmark credibility.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.