INQUIRING LINE

One AI wrote three research papers on its own and sent them to a real conference — only one passed review, and even that got pulled.

How do AI-generated papers perform when submitted to real conferences?

This explores what actually happens when fully AI-written research papers go through real peer review, and what those results tell us about the papers and about the review process judging them.


This explores what happens when papers written end to end by AI are sent to real peer review, and what the results say about both the papers and the reviewers. The best-documented test is small. Sakana AI's AI Scientist-v2 submitted three fully autonomous manuscripts to an ICLR 2025 workshop under double-blind review. One averaged 6.33 from reviewers, enough to clear the acceptance bar and rank in the top 45% of workshop submissions Can AI-generated papers pass peer review undetected?. The other two did not get through. The accepted paper was withdrawn under an agreement made in advance, and the authors later found a citation error in it. Their own verdict was that none of the three met main-conference standards Can AI systems generate research papers that pass peer review?. So the honest headline is that AI papers can get into a workshop now and then, but they aren't yet doing main-track science.

The less obvious point is that passing review may say as much about the reviewers as about the paper. Human review at AI conferences is already under strain. One position paper argues that authors, reviewers, and venues all share the blame, and points to measured biases such as review scores tracking how long a review is Can two-stage review and badges fix AI conference peer review?. AI reviewers are weaker still. They show a 'hivemind' effect, agreeing with each other more than humans do, and simply rewording a paper raises AI review scores by about 0.45 points with no change to the science Can AI systems safely replace human peer reviewers?. LLM judges also give higher scores to responses with fake references or polished formatting Can LLM judges be tricked without accessing their internals?. Some authors have already acted on this: 18 arXiv manuscripts were found with hidden instructions telling AI reviewers to rate them favorably Are hidden AI prompts in preprints a deceptive research practice?.

This loops back into how AI papers get made. The original AI Scientist's authors say the system only reaches its under-$15-per-paper cost because an automated reviewer scores each paper and feeds those scores back into idea generation Can automated review scale AI paper evaluation reliably?. Put that next to the gameability findings and a risk appears: a system tuned against an AI reviewer may learn to write papers that look good to reviewers rather than papers that are right. A finance demonstration shows how bad this can get. It produced 288 complete papers from 96 statistically significant signals, each with an invented theory written after the result and made-up citations Can AI generate hundreds of fake academic papers automatically?. Volume and plausibility are cheap, and review has trouble catching that.

The proposed fixes go in two directions. One builds verification into how papers are generated. Spark-to-Paper separates the model's judgment calls from deterministic checks that can actually be run, and requires the evidence plan to be set before any results are seen, which is a direct answer to the HARKing problem Can separating judgment from verification improve research paper reliability?. The other builds new venues. aiXiv is a dedicated venue for AI-generated research that uses repeated automated review-and-revise cycles with defenses against prompt injection, and reports measurable quality gains Can automated review loops handle AI-generated research at scale?. The open question these sources point to is less whether AI papers can pass review and more which review process is worth passing.


Sources 10 notes

Can AI-generated papers pass peer review undetected?

Sakana AI's end-to-end system produced a paper that scored 6.33 in double-blind ICLR 2025 workshop review, meeting acceptance thresholds, but was withdrawn under pre-agreed protocol. Authors later identified a citation error and judged none of three submissions suitable for main-track publication.

Can AI systems generate research papers that pass peer review?

AI Scientist-v2 submitted three fully autonomous manuscripts to ICLR; one averaged 6.33 from reviewers and ranked in the top 45% of workshop submissions. The authors acknowledged the work does not yet meet top-tier conference standards and withdrew the accepted paper before publication.

Can two-stage review and badges fix AI conference peer review?

Authors, reviewers, and venues all contribute to peer review failures at major AI conferences. A proposed two-stage system lets authors rate review quality before seeing verdicts, and a badge system rewards reviewer thoroughness, targeting measured biases like rating-length correlation.

Can AI systems safely replace human peer reviewers?

AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.

Can LLM judges be tricked without accessing their internals?

Research shows LLM evaluators systematically score higher when responses include fake references or rich formatting, independent of content quality. These biases are exploitable without model access, undermining AI benchmark credibility.

Show all 10 sources
Are hidden AI prompts in preprints a deceptive research practice?

Eighteen arXiv manuscripts contained concealed instructions directing AI reviewers to give positive assessments. The practice qualifies as questionable research conduct because concealment plus self-serving design violates ethics regardless of stated intent.

Can automated review scale AI paper evaluation reliably?

The AI Scientist's authors argue their system scales to sub-$15 per-paper cost only because they designed an automated reviewer. The reviewer's scores feed back into idea generation, allowing iterative research development at scale that manual review cannot match.

Can AI generate hundreds of fake academic papers automatically?

A demonstration showed LLMs generating 288 complete finance papers from 96 statistically significant signals, each with invented theoretical justifications and fabricated citations, proving academic HARKing can be automated at scale.

Can separating judgment from verification improve research paper reliability?

Spark-to-Paper architects paper generation as composable skills that isolate model judgment from executable, verifiable operations and require evidence specification before results are observed, reducing dependence on model correctness for consistency.

Can automated review loops handle AI-generated research at scale?

aiXiv demonstrates that iterative review-refine cycles with automated retrieval-augmented evaluation and prompt-injection defenses measurably enhance proposal and paper quality, addressing the structural gap where AI-generated research lacks appropriate publication venues.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.