Can AI systems safely replace human peer reviewers?
Explores whether AI reviewers meet two critical conditions for automation: maintaining diverse perspectives and resisting score manipulation. Tests whether current systems are ready to handle peer review at scale.
The position paper argues that "today's AI systems should not be used to produce paper reviews," and grounds this in two necessary conditions: C1, preservation of review diversity, and C2, resistance to gaming. On C1 it reports a hivemind effect: AI reviewers agree more within papers (IntraSim +8.7% to +9.8%) and across papers (InterSim +4.1% to +39.8%) than humans do, both in simulation and in real ICLR 2026 reviews. On C2 it introduces "paper laundering," a fully automated zero-shot LaTeX rewrite. Across 24 conditions (60 sampled ICLR 2026 papers, four prompts, two launderer models, three reviewer models), laundering raised AI review scores by +0.45 (p < 0.0001), at about $0.25 per paper. The paper takes its scale from a third party: Emi (2025) labels 15,899 of 75,800 ICLR 2026 reviews (21%) as AI-generated.
The paper's reason for holding AI to a higher bar is distributed versus centralized error. Human biases are spread across reviewers with different expertise and "partially cancel out" in aggregation. AI errors are correlated because models trained on similar data share biases, which the paper calls algorithmic monoculture. Gaming follows the same logic: gaming one human reviewer does not transfer to others, while "a single rewrite strategy can boost scores across models." The homogenization is measured, not only argued. Laundered papers become more similar to each other (pairwise similarity +6.5%, Cohen's d = 1.02), so the paper's worry is that a centralized system shapes "not only which papers are accepted but also how those papers are written."
The gaming result overlaps with How much does rhetorical style shift AI review scores?, which reaches a similar conclusion through a paired design, with the same paper in more and less favorable rhetorical versions. This excerpt adds the diversity condition and the cross-paper homogenization, which that note does not report. The paper also draws the opposite inference from Can human review keep pace with AI-accelerated research generation?: "the peer review crisis is real. An effective solution requires validated tools, not a simple replacement of human judgment." The ICML 2026 two-policy framework the excerpt cites is measured in Does banning LLM use in peer review change review outcomes?, but that study concerns reviewers' own LLM use, not AI systems writing reviews. The tested reviewers are general-purpose models and the agents of Bianchi et al. (2025b), not a purpose-built pipeline like Can inference scaling help reviewers catch errors humans miss?.
The excerpt does not establish several things. It gives no figures for the claim that AI-generated ratings are "less informative about final acceptance decisions" than human ratings, which appears only in the conclusion. The AI-review labels come from a third party, and the authors name a single reviewer prompt as a limitation, with the rest left to an appendix not included here. The claim that laundered scores reflect no "genuine improvement of scientific content" rests on the rewrites being purely textual; no human assessment of the laundered papers' substance appears in the excerpt. The narrower reading the evidence supports is that, for the models and prompt tested, AI reviewer scores move with textual rewrites, so a score cannot be read as merit without a robustness check. The broader case for accountability and democratic legitimacy is argued, not measured, and the paper's own answer is "cautious automation" until those risks are studied.
Inquiring lines that read this note 77
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can AI systems perform peer review as effectively as humans?- What specific errors did participants report finding in the AI-generated reviews?
- Can multi-stage AI review pipelines catch scientific flaws better than simple language models?
- Did adding AI reviews actually change peer review decisions or paper outcomes?
- Should rhetorical polish in AI reviews be separated from actual technical accuracy?
- Can human reviewers detect when papers have been rewritten by AI?
- How much of ICLR 2026 peer review was already conducted by AI?
- Do AI reviews depend more on writing style than scientific merit?
- Should AI research papers require dedicated automated review systems instead of existing journals?
- Can workshop acceptance rates reliably measure AI research quality compared to main conferences?
- Does review length bias affect acceptance decisions at major conferences?
- How much does reviewer consistency vary across different papers at NeurIPS?
- How often do researchers violate rules about AI use in review?
- Can automated review systems catch deep methodological flaws or only surface issues?
- What specific tasks do reviewers use AI for most often?
- Could AI improve peer review rigor and catch human-missed errors?
- How did researchers measure whether GPT-4 and human reviewers identified the same issues?
- Does matching reviewer points actually mean the feedback is accurate or correct?
- Why did rejected papers show higher overlap between GPT-4 and human reviewers?
- Could AI feedback work as a substitute for human peer review entirely?
- Can polished AI text fool both reviewers and detection methods?
- Can computational inference scaling catch flaws that human expert reviewers miss?
- How often do false positives from detection tools actually occur in peer review?
- What effects do preprint servers have on scientific consensus formation?
- How can arXiv and journals scale quality control for AI-generated research?
- Do AI-generated research reviews score papers higher than human reviewers do?
- How often do researchers suspect peer reviews are written by AI?
- Does AI content in reviews correlate with differences in paper quality control?
- Can human reviewers reliably detect AI-written peer review text by sight?
- Can machine review catch flaws in AI-generated work that humans miss?
- How do automated reviewers detect flaws that human experts miss in manuscripts?
- Should AI-generated papers use specialized review venues instead of traditional journals?
- Why do peer reviewers favor novel ideas that later fail in execution?
- Do peer reviewers actually follow restrictions on using AI tools themselves?
- Are refereed venues also overwhelmed by AI-generated low-quality submissions?
- Can automated systems scale peer review faster than human moderators?
- Can AI reviewers detect deep theoretical flaws that human experts miss?
- Does rhetorical presentation bias reviewers against substantive scientific contributions?
- Can agentic AI systems catch flaws in manuscripts that human reviewers consistently miss?
- Why do individual peer reviewers show such low agreement on research merit?
- Could hidden prompts be inserted during review and removed before publication?
- How fast is scientific publishing growing relative to reviewer capacity?
- Could automated review systems handle AI-generated research at scale?
- What role do conference organizers play in accepting problematic articles?
- Which feedback loops in AI-mediated review remain unmeasured or rarely observed directly?
- Do shortened peer review timelines correlate with lower quality publications?
- Does an automated reviewer's output actually match human review accuracy?
- How do AI-generated papers perform when submitted to real conferences?
- What limitations did the authors acknowledge about their automated reviewer?
- Can feeding review scores back into idea generation improve research quality?
- Why do researchers resist using AI for peer review specifically?
- Can AI systems write and review research while operating outside traditional PDF constraints?
- How should hiring and promotion weigh AI-inflated research output?
- Can automated reviewers actually handle the review load AI creates?
- How does opaque AI methodology undermine peer review and reproducibility?
- Can technical accuracy in AI training data replace human review before publication?
- Can traditional complexity measures still signal research quality in AI-era papers?
- What role should humans play in reviewing and approving AI-generated research?
- Are paper mills using NHANES data to automate single-factor research?
- What would a practical reviewer checklist for autonomous research systems need to include?
- Can AI agents themselves become reliable reviewers of other autonomous research systems?
- Why do current AI systems struggle with researcher judgment and taste?
- Can LLM-generated reference reviews detect machine-written peer review submissions?
- What quality differences exist between flagged and unflagged peer reviews?
- What mechanisms drive rating compression in fully LLM-generated peer reviews?
- What prevents venues from implementing two-way feedback systems at scale?
- Do reviewer reward badges risk encouraging lenient or superficial reviews?
- Do stricter AI policies actually change how reviewers score manuscripts?
- Do peer review policies banning LLM use actually change reviewer behavior and decisions?
- Why does author withdrawal authority matter for preprint accountability?
- Do conference policies banning LLM use actually reduce AI involvement in reviews?
- What happens when reviewers use AI tools against journal policy?
- Do metareviewers and regular reviewers use LLMs differently in peer review?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
How much does rhetorical style shift AI review scores?
When manuscripts are rewritten to improve rhetoric while keeping scientific content identical, do LLM reviewers change their scores? Understanding this matters for ensuring AI-assisted peer review evaluates substance, not polish.
paired-rewrite study reaching the same gaming finding; this excerpt adds the diversity condition
-
Does banning LLM use in peer review change review outcomes?
Can policies restricting or allowing AI tools shift how reviewers score papers and make decisions? This matters because review quality and fairness depend on consistent standards.
measures the ICML two-policy framework this excerpt cites; about reviewers' own LLM use, not AI-written reviews
-
Can human review keep pace with AI-accelerated research generation?
As AI systems generate hypotheses, code, and proofs faster than humans can verify them, does the bottleneck at peer review force verification itself to become automated? What governance structures enable this transition safely?
the opposite inference: the paper says the crisis calls for validated tools, not replacement
-
Can inference scaling help reviewers catch errors humans miss?
Explores whether spending extra compute at review time—checking proofs and experiments line by line—can surface deep flaws that evade human expert reviewers, and how this scales with AI-assisted submissions.
purpose-built reviewer this excerpt does not test; the gaming result leaves its robustness open
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Stop Automating Peer Review Without Rigorous Evaluation
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?
- How to Find Fantastic AI Papers: Self-Rankings as a Powerful Predictor of Scientific Impact Beyond Peer Review
- Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards
- Towards End-to-End Automation of AI Research
Original note title
AI reviewers fail both necessary conditions for peer review automation — a hivemind effect and scores trivially gameable through rewriting