INQUIRING LINE

Does AI-assisted peer review catch flawed papers more often, or let them slip through, and has anyone measured it?

Does AI content in reviews correlate with differences in paper quality control?

This explores whether peer reviews that are partly or wholly written by AI go along with stronger or weaker quality control on the papers being reviewed, meaning whether flawed work gets caught or slips through.


This explores whether AI involvement in peer review goes along with better or worse quality control on the papers being judged. The direct answer: the collection has no study that measures how much AI-written text appears in reviews and then links that share to the quality of accepted papers. What it does have is evidence about what AI does to each step of review. That evidence points in two opposite directions, depending on how the AI is used.

In the hopeful direction, AI used as a checker seems to catch errors people miss. An agentic reviewer that spends extra compute working through proofs and experiments line by line found serious flaws in papers that had already passed human review at STOC and ICML Can inference scaling help reviewers catch errors humans miss?. In a randomized trial at ICLR 2025, AI feedback on reviewers' drafts led 27% of them to revise, and blinded raters judged the revised reviews more specific and informative Can LLM feedback help peer reviewers improve their own reviews?. At AAAI-26, survey respondents said they preferred the labeled AI review over the human ones on technical accuracy. The organizers didn't report how large that preference was Did conference reviewers prefer AI reviews over human ones?. Closed loops where an AI reviews a paper and the paper is then revised also measurably improved AI-generated research Can automated review loops handle AI-generated research at scale?.

In the worrying direction, AI used as the judge creates weaknesses that human review doesn't have. AI reviewers agree with each other more than human reviewers do, so the independent viewpoints that make review useful start to disappear. A simple automated rewrite of a paper's wording raised AI scores by 0.45 points without changing the science Can AI systems safely replace human peer reviewers?. Authors have already noticed: 18 arXiv manuscripts were found with hidden instructions telling AI reviewers to be positive Are hidden AI prompts in preprints a deceptive research practice?. On the submission side, one fully AI-generated paper reached a passing score at an ICLR workshop, though its own authors later found a citation error and judged none of their three submissions good enough for the main conference Can AI-generated papers pass peer review undetected?.

Here is the part you may not have expected to care about: the correlation the question asks for is hard to measure at all. People identify AI-generated content at about chance level Can people reliably spot content made by AI?. A writing study found people edited AI text only 23% of the time, and their edits left it about 96% the same Do writers actually edit AI-generated text before publishing?. If reviewers behave the same way, AI wording passes through largely untouched and can't be easily spotted. A survey of 230 publications describes production, review, manipulation and defense as one connected arms race, where each side adapts to the other. It finds the evidence strongest for the early stages and weakest for long-term effects on the research ecosystem, which is where a quality-control correlation would show up Does AI create a coupled arms race in research production and review?.

The pattern across the collection is that how AI is used matters more than how much. AI that checks claims or feeds back into human judgment tends to improve quality control. AI that replaces human judgment adds new failure modes: reviews that all agree, scores that can be gamed, and hidden prompts. That is why some proposals focus on accountability rather than detection. They spread responsibility across authors, reviewers and venues, and let authors rate a review before seeing the decision Can two-stage review and badges fix AI conference peer review?. A paper-generation design that keeps model judgment separate from automatic checks follows the same logic from the writing side Can separating judgment from verification improve research paper reliability?.


Sources 12 notes

Can inference scaling help reviewers catch errors humans miss?

PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.

Can LLM feedback help peer reviewers improve their own reviews?

A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.

Did conference reviewers prefer AI reviews over human ones?

At AAAI-26, every full-review paper received one labeled AI review generated by a multi-stage pipeline. Survey respondents reported preferring these AI reviews over human reviews on dimensions like technical accuracy, though the preference size and statistical strength were not disclosed.

Can automated review loops handle AI-generated research at scale?

aiXiv demonstrates that iterative review-refine cycles with automated retrieval-augmented evaluation and prompt-injection defenses measurably enhance proposal and paper quality, addressing the structural gap where AI-generated research lacks appropriate publication venues.

Can AI systems safely replace human peer reviewers?

AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.

Show all 12 sources
Are hidden AI prompts in preprints a deceptive research practice?

Eighteen arXiv manuscripts contained concealed instructions directing AI reviewers to give positive assessments. The practice qualifies as questionable research conduct because concealment plus self-serving design violates ethics regardless of stated intent.

Can AI-generated papers pass peer review undetected?

Sakana AI's end-to-end system produced a paper that scored 6.33 in double-blind ICLR 2025 workshop review, meeting acceptance thresholds, but was withdrawn under pre-agreed protocol. Authors later identified a citation error and judged none of three submissions suitable for main-track publication.

Can people reliably spot content made by AI?

A 30-study systematic review found that humans cannot reliably distinguish AI-generated from human-created content across text, image, and voice modalities. Accuracy generally clusters around chance and has not kept pace with improvements in AI realism.

Do writers actually edit AI-generated text before publishing?

Writers edited AI-generated paragraphs only 23% of the time, with edits averaging 96% similarity to the original. This means AI's opinionated and distorted voice propagates with minimal human filtering before publication.

Does AI create a coupled arms race in research production and review?

A survey of 230 publications reveals production scaling, evaluation automation, manipulation, defenses, evasion, and ecosystem feedback as linked response relations among actors. Evidence is strongest for early stages and weakens toward long-horizon adaptation and feedback.

Can two-stage review and badges fix AI conference peer review?

Authors, reviewers, and venues all contribute to peer review failures at major AI conferences. A proposed two-stage system lets authors rate review quality before seeing verdicts, and a badge system rewards reviewer thoroughness, targeting measured biases like rating-length correlation.

Can separating judgment from verification improve research paper reliability?

Spark-to-Paper architects paper generation as composable skills that isolate model judgment from executable, verifiable operations and require evidence specification before results are observed, reducing dependence on model correctness for consistency.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.