INQUIRING LINE

Sources show no measured rate for how often AI detectors wrongly flag peer reviews, so venues treat flags as one signal, not a verdict.

How often do false positives from detection tools actually occur in peer review?

This explores how often AI-text detectors wrongly flag human-written reviews or papers as machine-generated, and what that error rate means for conferences that use these tools.


This explores how often AI-detection tools wrongly accuse human-written reviews or papers of being machine-generated, and how much that matters in peer review. The short answer is that the corpus does not contain a measured false-positive rate for peer review. What it does show is more useful: the venues that rely on these detectors behave as if the error rate is high enough that a flag can't be treated as a verdict.

The clearest case is ICLR 2026. Its program chairs did not let detector output reject anything on its own. Flags went to area chairs as one quality signal among several, and several human review steps sat between a flag and any penalty, precisely to absorb false positives How can conferences detect and handle LLM misuse in peer review?. The chairs used automatic enforcement in only one place: hallucinated references. A citation either exists or it doesn't, so it can be checked directly. That suggests a principle you can apply elsewhere. Statistical guesses about style get human judgment, and checkable facts get automatic enforcement.

The headline numbers about AI in review come from the same kind of detectors, so they carry the same uncertainty. Pangram Labs estimated that 21% of ICLR reviews were fully AI-generated and that more than half had some AI involvement How much AI content appears in peer review at ICLR?. Across a whole corpus, a population estimate like that can hold up even when individual calls are wrong. Applied to one reviewer's report, it can't. The more interesting finding sits next to the estimate: reviews with more AI text gave higher scores. That pattern doesn't depend on any single flag being right, and it points to a cost of *missed* detections that gets less attention than false accusations.

The detection problem also runs in the opposite direction. One fully AI-generated paper from Sakana AI cleared double-blind workshop review at ICLR 2025, and it was caught only because the authors had agreed in advance to withdraw it Can AI-generated papers pass peer review undetected?. Eighteen arXiv manuscripts hid prompts telling AI reviewers to be positive Are hidden AI prompts in preprints a deceptive research practice?. Simply rewriting a paper's text can raise AI review scores by almost half a point without changing the science Can AI systems safely replace human peer reviewers?. So detectors face misses and false alarms together, and any threshold trades one for the other.

If you want a framework for that trade, a security-monitoring paper outside peer review proposes comparing detection designs at equal review cost and equal false-alert workload Does added monitoring improve protection at acceptable cost?. It reports no results yet, but the framing transfers well. The useful question is less "how often are detectors wrong?" than "how much human attention does each false alarm cost, and is there any to spare?" In peer review there often isn't, because overloaded reviewers are already lowering accuracy across the system Does peer review quality collapse under submission overload?.


Sources 7 notes

How can conferences detect and handle LLM misuse in peer review?

Program chairs used imperfect detectors as one input for area chairs rather than automated filters, but desk-rejected papers with confirmed fabricated references as a tractable enforcement point. Multiple human review steps mitigated false positives.

How much AI content appears in peer review at ICLR?

Pangram Labs' analysis of ICLR's public review corpus estimates 21% of reviews were fully AI-generated and over half had some AI involvement. Reviews with more AI text received systematically higher scores, suggesting AI may amplify positive bias rather than just rephrase human judgment.

Can AI-generated papers pass peer review undetected?

Sakana AI's end-to-end system produced a paper that scored 6.33 in double-blind ICLR 2025 workshop review, meeting acceptance thresholds, but was withdrawn under pre-agreed protocol. Authors later identified a citation error and judged none of three submissions suitable for main-track publication.

Are hidden AI prompts in preprints a deceptive research practice?

Eighteen arXiv manuscripts contained concealed instructions directing AI reviewers to give positive assessments. The practice qualifies as questionable research conduct because concealment plus self-serving design violates ethics regardless of stated intent.

Can AI systems safely replace human peer reviewers?

AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.

Show all 7 sources
Does added monitoring improve protection at acceptable cost?

The paper designs a controlled comparison of isolated actions, rolling windows, known groups, and prospectively discovered episodes at equal review cost and false-alert workload, but the excerpt provides no empirical results showing whether added monitoring improves protection.

Does peer review quality collapse under submission overload?

A two-journal model shows that rising submissions overtax unpaid reviewers, forcing journals to recruit less qualified reviewers or overload existing ones, which drops review accuracy and incentivizes authors to submit more speculatively, driving submissions higher. The mechanism is structural but its empirical strength remains to be measured.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.