INQUIRING LINE

Peer-reviewed journals and conferences face similar pressure, and the review step is being automated and gamed too.

Are refereed venues also overwhelmed by AI-generated low-quality submissions?

This explores whether peer-reviewed conferences and journals are being flooded with low-quality AI-generated papers, the way open platforms are, and what the corpus says about how review is holding up under that pressure.


This explores whether refereed venues, meaning conferences and journals with peer review, are being swamped by low-quality AI-generated papers. The direct answer is that the corpus doesn't measure a flood. None of these notes counts how many AI-written submissions reach a venue or tracks how that number changes. What the corpus does show is more useful: the pressure is real, and the review side is being automated and gamed at the same time as the writing side. The best framing comes from a survey of 230 publications, which describes an arms race of six linked dynamics. More papers are produced, review is automated, people manipulate the review, venues build defenses, people evade those defenses, and the whole system feeds back on itself Does AI create a coupled arms race in research production and review?. The evidence is strong for the early stages, like more papers being produced, and weak for the long-run effects, which is exactly where the 'overwhelmed' question sits.

The most concrete test case is a stress test, not a flood. Sakana's AI Scientist-v2 sent three fully AI-generated manuscripts to an ICLR 2025 workshop. One averaged a reviewer score of 6.33, enough to be accepted, before it was withdrawn under a protocol agreed in advance Can AI-generated papers pass peer review undetected?. The authors' own verdict is the more revealing part. They later found a citation error and judged that none of the three papers met main-conference standards Can AI systems generate research papers that pass peer review?. So the gate leaks at the workshop level, where review is lighter, but no slop wave has been documented getting through.

The less obvious risk is that AI might overwhelm venues through their reviewers rather than their submission piles. When conferences lean on AI to handle the volume, those reviewers show a 'hivemind' effect: they agree with each other more than humans do. They are also easy to game. Rewriting a paper's text, with no change to the science, raised AI review scores by 0.45 points Can AI systems safely replace human peer reviewers?. Research on AI judges more broadly finds the same weak spots. Fake references ('authority bias') and polished formatting ('beauty bias') raise scores without any access to the model Can LLM judges be fooled by fake credentials and formatting?. A low-quality paper doesn't have to fool a tired human if it can fool the automated reviewer that triage depends on.

The corpus also points to defenses that look less like filtering and more like strengthening review. At ICLR 2025, a randomized trial gave reviewers optional AI feedback on their reviews, and 27% of them revised to be more specific Can LLM feedback help peer reviewers improve their own reviews?. PAT, an agentic reviewer that spends extra compute checking proofs line by line, found serious flaws in papers that had already passed human review at STOC and ICML Can inference scaling help reviewers catch errors humans miss?. One position paper argues the failures are shared among authors, reviewers, and venues. It proposes letting authors rate a review's quality before they see the verdict, and giving reviewers badges for thoroughness Can two-stage review and badges fix AI conference peer review?. Another approach builds a separate venue, aiXiv, for AI-generated research, where automated review-and-revise loops raise quality before anything is published Can automated review loops handle AI-generated research at scale?.

The takeaway you might not expect: the documented weak point of refereed venues isn't the number of submissions. It's the cheapness of the signals reviewers use to judge quality. If formatting, citations, and fluent prose can raise scores, then AI 'slop' doesn't need to be good, only well dressed. Whether venues are actually being overwhelmed is still an open, unmeasured question in this collection.


Sources 9 notes

Does AI create a coupled arms race in research production and review?

A survey of 230 publications reveals production scaling, evaluation automation, manipulation, defenses, evasion, and ecosystem feedback as linked response relations among actors. Evidence is strongest for early stages and weakens toward long-horizon adaptation and feedback.

Can AI-generated papers pass peer review undetected?

Sakana AI's end-to-end system produced a paper that scored 6.33 in double-blind ICLR 2025 workshop review, meeting acceptance thresholds, but was withdrawn under pre-agreed protocol. Authors later identified a citation error and judged none of three submissions suitable for main-track publication.

Can AI systems generate research papers that pass peer review?

AI Scientist-v2 submitted three fully autonomous manuscripts to ICLR; one averaged 6.33 from reviewers and ranked in the top 45% of workshop submissions. The authors acknowledged the work does not yet meet top-tier conference standards and withdrew the accepted paper before publication.

Can AI systems safely replace human peer reviewers?

AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.

Can LLM judges be fooled by fake credentials and formatting?

Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.

Show all 9 sources
Can LLM feedback help peer reviewers improve their own reviews?

A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.

Can inference scaling help reviewers catch errors humans miss?

PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.

Can two-stage review and badges fix AI conference peer review?

Authors, reviewers, and venues all contribute to peer review failures at major AI conferences. A proposed two-stage system lets authors rate review quality before seeing verdicts, and a badge system rewards reviewer thoroughness, targeting measured biases like rating-length correlation.

Can automated review loops handle AI-generated research at scale?

aiXiv demonstrates that iterative review-refine cycles with automated retrieval-augmented evaluation and prompt-injection defenses measurably enhance proposal and paper quality, addressing the structural gap where AI-generated research lacks appropriate publication venues.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.