INQUIRING LINE

Sakana's AI Scientist uses an AI reviewer to score its own papers cheaply — but how much should anyone trust that reviewer's judgment?

What limitations did the authors acknowledge about their automated reviewer?

This reads the question as asking about The AI Scientist (Sakana AI's system for automating research end to end): what its authors admitted their AI paper reviewer can't do, and what the rest of the corpus says about where automated reviewers fall short.


This reads the question as asking about The AI Scientist and the automated reviewer its authors built to score the papers it writes. One point first: the corpus records why the authors needed that reviewer, but it does not record their own list of its limitations. The reviewer is what lets the system run at under $15 per paper. Its scores also feed back into choosing which ideas to pursue next, so the whole loop depends on how much you can trust it (Can automated review scale AI paper evaluation reliably?). If you want the authors' own caveats about the reviewer, you'll need the paper itself.

The authors were open about a nearby gap: their reviewer's verdict was not good enough to stand alone. When they tested the follow-up system, they sent three fully AI-written papers to human reviewers at an ICLR 2025 workshop. One scored 6.33, which met the acceptance bar. Even so, the authors judged that none of the three was good enough for the main conference. They later found a citation error in the paper that was accepted, and they withdrew it under a protocol they had agreed to in advance (Can AI systems generate research papers that pass peer review?, Can AI-generated papers pass peer review undetected?). In effect, they used human peer review, plus their own checking afterwards, as the final test of quality. A citation error that got past both the system and human reviewers is the kind of problem an automated reviewer is supposed to catch.

Other work in the corpus spells out the risks that the AI Scientist authors' optimism skips over. AI reviewers show a 'hivemind' effect: they agree with each other more than human reviewers do, so running several of them doesn't give you independent opinions. They are also easy to game. Rewording a paper, with no change to the science, raised AI review scores by 0.45 points (Can AI systems safely replace human peer reviewers?). This is especially awkward for The AI Scientist, because its reviewer's scores steer what the system writes next. A paper writer guided by those scores could learn to please the reviewer rather than do better science. Deliberate attacks have already appeared in the wild: eighteen arXiv preprints contained hidden instructions telling AI reviewers to give positive reviews (Are hidden AI prompts in preprints a deceptive research practice?). The aiXiv venue for AI-generated papers had to build defenses against these 'prompt injections' into its review loop (Can automated review loops handle AI-generated research at scale?).

The less obvious lesson is that a single overall score from an AI reviewer is the weak point. Approaches that work better break review into smaller pieces that can be checked. PAT, an AI reviewer that spends extra computing time checking proofs and experiments line by line, catches 34% more math errors than an AI reviewer working in a single pass. It has even found serious flaws in papers that had passed human review at STOC and ICML (Can inference scaling help reviewers catch errors humans miss?). Spark-to-Paper takes a similar line on the writing side. It keeps the model's judgment calls separate from steps a program can verify, so that less depends on whether the model gets things right (Can separating judgment from verification improve research paper reliability?). The corpus suggests the fix is not a smarter reviewer that hands out grades. It is a reviewer that shows its work, one claim at a time.


Sources 8 notes

Can automated review scale AI paper evaluation reliably?

The AI Scientist's authors argue their system scales to sub-$15 per-paper cost only because they designed an automated reviewer. The reviewer's scores feed back into idea generation, allowing iterative research development at scale that manual review cannot match.

Can AI systems generate research papers that pass peer review?

AI Scientist-v2 submitted three fully autonomous manuscripts to ICLR; one averaged 6.33 from reviewers and ranked in the top 45% of workshop submissions. The authors acknowledged the work does not yet meet top-tier conference standards and withdrew the accepted paper before publication.

Can AI-generated papers pass peer review undetected?

Sakana AI's end-to-end system produced a paper that scored 6.33 in double-blind ICLR 2025 workshop review, meeting acceptance thresholds, but was withdrawn under pre-agreed protocol. Authors later identified a citation error and judged none of three submissions suitable for main-track publication.

Can AI systems safely replace human peer reviewers?

AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.

Are hidden AI prompts in preprints a deceptive research practice?

Eighteen arXiv manuscripts contained concealed instructions directing AI reviewers to give positive assessments. The practice qualifies as questionable research conduct because concealment plus self-serving design violates ethics regardless of stated intent.

Show all 8 sources
Can automated review loops handle AI-generated research at scale?

aiXiv demonstrates that iterative review-refine cycles with automated retrieval-augmented evaluation and prompt-injection defenses measurably enhance proposal and paper quality, addressing the structural gap where AI-generated research lacks appropriate publication venues.

Can inference scaling help reviewers catch errors humans miss?

PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.

Can separating judgment from verification improve research paper reliability?

Spark-to-Paper architects paper generation as composable skills that isolate model judgment from executable, verifiable operations and require evidence specification before results are observed, reducing dependence on model correctness for consistency.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.