INQUIRING LINE

When AI can produce more papers than people can review, should AI reviewers judge that research instead?

Should AI research papers require dedicated automated review systems instead of existing journals?

This explores whether research written by AI systems (not research about AI) should go to its own publishing venue judged by automated reviewers, rather than through the usual human peer review at journals and conferences.


This explores whether AI-written research should get its own publishing venue with automated reviewers, instead of going through the human peer review that journals and conferences use now. The case for a separate venue is real. AI systems can already produce papers that pass for workshop-level work. The original AI Scientist ran the whole cycle from idea to self-reviewed manuscript on its own Can one AI system complete a full research cycle end-to-end?. Its successor submitted three fully AI-generated papers to an ICLR workshop, and one scored 6.33 under double-blind review, which was above the acceptance threshold Can AI-generated papers pass peer review undetected?. If machines can produce papers at that volume, human reviewers can't keep up. aiXiv's answer is a dedicated venue where automated reviewers critique a paper, the AI author revises it, and the cycle repeats. It reports that this loop measurably improves the papers Can automated review loops handle AI-generated research at scale?.

The catch is that the reviewers in such a venue would be AI too, and the corpus has a pointed warning about that. AI reviewers show a 'hivemind' effect: they agree with each other more than human reviewers do, so ten AI reviews can amount to one opinion repeated ten times. They are also easy to game. Rewriting a paper's text with no change to the science raised AI review scores by about 0.45 points Can AI systems safely replace human peer reviewers?. People are already trying this. Hidden instructions telling AI reviewers to praise the paper turned up in 18 arXiv manuscripts Are hidden AI prompts in preprints a deceptive research practice?. A survey of 230 publications describes the situation as an arms race. AI makes papers cheaper to produce, so review gets automated. Then people learn to manipulate the automated reviewers, defenses follow, and people find ways around the defenses Does AI create a coupled arms race in research production and review?. A venue where AI writes and AI reviews puts both sides of that race inside one system.

The more surprising finding is where AI review is clearly working: as a tool that helps human reviewers. At ICLR 2025, a randomized trial gave reviewers optional feedback from Claude-based agents. 27 percent of them revised their reviews, and independent raters judged the revised reviews more informative Can LLM feedback help peer reviewers improve their own reviews?. An agentic reviewer called PAT uses extra compute to check proofs and experiments line by line, and it found serious errors in papers already accepted at STOC and ICML Can inference scaling help reviewers catch errors humans miss?. In both cases AI handles narrow, checkable work like math and specificity, and judgment stays with people. Spark-to-Paper applies the same split to the writing side. It keeps the model's judgment separate from checks that can be run and verified, so the paper's reliability depends less on the model being right Can separating judgment from verification improve research paper reliability?.

The corpus doesn't support choosing between 'automated venue' and 'existing journals.' It supports a hybrid in both kinds of venue: machines do verification, humans do judgment, and the incentives get repaired. Some of peer review's problems have nothing to do with AI. One position paper argues that authors, reviewers, and venues all share the blame. It proposes letting authors rate review quality before they see the decision, and rewarding thorough reviewers Can two-stage review and badges fix AI conference peer review?. Automating a broken incentive system would just make it break faster. The corpus has no long-term evidence on how a fully separate AI venue would develop over time, and the arms-race survey says its evidence is weakest on exactly that question.


Sources 10 notes

Can one AI system complete a full research cycle end-to-end?

The AI Scientist performed ideation, coding, experiments, writing, and self-review autonomously, producing a manuscript that passed the first round at a machine learning workshop with 70% acceptance rate. Five ensemble reviewers and an area-chair model judged the output against NeurIPS guidelines.

Can AI-generated papers pass peer review undetected?

Sakana AI's end-to-end system produced a paper that scored 6.33 in double-blind ICLR 2025 workshop review, meeting acceptance thresholds, but was withdrawn under pre-agreed protocol. Authors later identified a citation error and judged none of three submissions suitable for main-track publication.

Can automated review loops handle AI-generated research at scale?

aiXiv demonstrates that iterative review-refine cycles with automated retrieval-augmented evaluation and prompt-injection defenses measurably enhance proposal and paper quality, addressing the structural gap where AI-generated research lacks appropriate publication venues.

Can AI systems safely replace human peer reviewers?

AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.

Are hidden AI prompts in preprints a deceptive research practice?

Eighteen arXiv manuscripts contained concealed instructions directing AI reviewers to give positive assessments. The practice qualifies as questionable research conduct because concealment plus self-serving design violates ethics regardless of stated intent.

Show all 10 sources
Does AI create a coupled arms race in research production and review?

A survey of 230 publications reveals production scaling, evaluation automation, manipulation, defenses, evasion, and ecosystem feedback as linked response relations among actors. Evidence is strongest for early stages and weakens toward long-horizon adaptation and feedback.

Can LLM feedback help peer reviewers improve their own reviews?

A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.

Can inference scaling help reviewers catch errors humans miss?

PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.

Can separating judgment from verification improve research paper reliability?

Spark-to-Paper architects paper generation as composable skills that isolate model judgment from executable, verifiable operations and require evidence specification before results are observed, reducing dependence on model correctness for consistency.

Can two-stage review and badges fix AI conference peer review?

Authors, reviewers, and venues all contribute to peer review failures at major AI conferences. A proposed two-stage system lets authors rate review quality before seeing verdicts, and a badge system rewards reviewer thoroughness, targeting measured biases like rating-length correlation.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.