AI now writes papers and reviews them, but do we measure what happens when each side reacts to the other?
Which feedback loops in AI-mediated review remain unmeasured or rarely observed directly?
This explores which cause-and-effect cycles in AI-assisted peer review, where AI writes papers, AI reviews them and authors respond to AI reviewers, have been claimed or assumed but not directly measured.
This explores which back-and-forth cycles in AI-assisted peer review have actually been measured, and which are only assumed. The clearest map comes from a survey of 230 publications. It describes research production and review as a linked arms race with six parts: more papers, automated evaluation, manipulation, defenses, evasion, and feedback across the whole system. The survey finds the strongest evidence at the start of that chain. The evidence gets weaker for the parts that take time, such as people adapting and those changes feeding back into the system Does AI create a coupled arms race in research production and review?. The corpus has a solid record of first moves. It has much less on what happens after those moves reach the people and systems on the other side.
The first missing measurement is the step from 'possible' to 'actually happening.' We know AI reviewers can be gamed easily: rewriting a paper's text with no new science raised AI scores by 0.45 points, and AI reviewers agree with each other more than human reviewers do Can AI systems safely replace human peer reviewers?. That is a lab demonstration. Nobody has shown how many authors now polish their papers for machine readers, or whether reviewer models change in response. The AI-authored papers are similar. Sakana's AI Scientist-v2 got one of three manuscripts past an ICLR workshop review, but the paper was withdrawn under an agreement made in advance Can AI-generated papers pass peer review undetected? Can AI systems generate research papers that pass peer review?. Because of that design, nobody saw what happens after an AI paper is accepted: whether it gets cited, built on, or found to be wrong.
The second gap is follow-through. At ICLR 2025, 27% of reviewers who received LLM feedback revised their reviews, and independent raters judged the revisions more informative Can LLM feedback help peer reviewers improve their own reviews?. The trial measured the review itself. It did not show whether better reviews led authors to improve their papers or changed final decisions. In ICML 2026's randomized test, it made almost no difference to scores whether reviewers were banned from using LLMs or allowed limited use. Large shares of reviewers also broke whichever rule they were given Does banning LLM use in peer review change review outcomes?. So the policy can't be judged on its own terms: the conference didn't fully control how much AI went into the reviews. The fix proposed in one position paper is to let authors rate reviews before seeing the verdict and give reviewers badges for thoroughness Can two-stage review and badges fix AI conference peer review?. That is a feedback loop that so far exists only on paper.
The third gap is a loop that may run almost invisibly, through writing style. Writers edited AI-drafted paragraphs only 23% of the time, and their edits left the text about 96% the same Do writers actually edit AI-generated text before publishing?. AI assistance also shifted how readers saw the writer on all 29 traits tested, toward sounding more confident and more extreme Does AI writing assistance change how readers perceive the writer?. Those studies weren't about peer review. Still, they suggest that AI-polished reviews and AI-polished papers could change how authors and reviewers come across to each other without anyone noticing. No study in this collection has measured that inside review.
The loops that do get measured are mostly self-contained ones. aiXiv shows that cycles of automated review and revision improve quality inside its own platform Can automated review loops handle AI-generated research at scale?. The PAT reviewer caught math errors that had passed human review at STOC and ICML Can inference scaling help reviewers catch errors humans miss?. That result points to an uncomfortable unknown: how many such flaws human review lets through in general. Agent-based judges have their own hidden loop too. A memory module passed errors from one step to the next, and those errors built up inside an evaluator that otherwise looked very stable Can agents evaluate AI outputs more reliably than language models?. Overall, the collection is good at measuring single moves and closed systems. The loops that run between people, between conferences and across years are where the evidence thins out.
Sources 12 notes
A survey of 230 publications reveals production scaling, evaluation automation, manipulation, defenses, evasion, and ecosystem feedback as linked response relations among actors. Evidence is strongest for early stages and weakens toward long-horizon adaptation and feedback.
AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.
Sakana AI's end-to-end system produced a paper that scored 6.33 in double-blind ICLR 2025 workshop review, meeting acceptance thresholds, but was withdrawn under pre-agreed protocol. Authors later identified a citation error and judged none of three submissions suitable for main-track publication.
AI Scientist-v2 submitted three fully autonomous manuscripts to ICLR; one averaged 6.33 from reviewers and ranked in the top 45% of workshop submissions. The authors acknowledged the work does not yet meet top-tier conference standards and withdrew the accepted paper before publication.
A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.
Show all 12 sources
A randomized experiment at ICML 2026 found that prohibiting LLM use versus allowing limited use barely changed paper scores, decisions, or reviewer confidence. Meanwhile, substantial fractions of reviewers broke whichever rule they were given.
Authors, reviewers, and venues all contribute to peer review failures at major AI conferences. A proposed two-stage system lets authors rate review quality before seeing verdicts, and a badge system rewards reviewer thoroughness, targeting measured biases like rating-length correlation.
Writers edited AI-generated paragraphs only 23% of the time, with edits averaging 96% similarity to the original. This means AI's opinionated and distorted voice propagates with minimal human filtering before publication.
A study of 2,939 writers and 11,091 readers found AI assistance shifted every tested dimension—29 total—toward extremism, confidence, quality, agreeableness, and perceived privilege. Distortions were statistically significant and directional, not random noise.
aiXiv demonstrates that iterative review-refine cycles with automated retrieval-augmented evaluation and prompt-injection defenses measurably enhance proposal and paper quality, addressing the structural gap where AI-generated research lacks appropriate publication venues.
PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.
Eight-module agentic evaluation achieved 0.27% judge shift versus 31% for LLM-as-a-Judge on complex tasks. However, the memory module cascaded errors, revealing that agentic systems need error isolation mechanisms to maintain gains.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stop Automating Peer Review Without Rigorous Evaluation
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?
- The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing
- Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards
- AI for Auto-Research: Roadmap & User Guide
- Pangram Predicts 21% of ICLR Reviews are AI-Generated