Can LLM feedback help peer reviewers improve their own reviews?
This randomized trial tested whether optional AI-generated suggestions on review quality would prompt reviewers to revise, and whether those revisions would be more useful to authors and decision-makers.
The paper's central claim is that optional LLM feedback on a reviewer's own review, gated by automated reliability tests, made reviews more specific and led many reviewers to revise them. The Review Feedback Agent, five LLMs with Claude Sonnet 3.5 as the backbone, ran at ICLR 2025 as a randomized control trial. The authors report that 26.6% of reviewers who received feedback updated their reviews (27% in the abstract), incorporating 12,222 suggestions. Blinded ML researchers rated the revised reviews "more informative and clearer," and reviewers who updated lengthened them by an average of 80 words. Rebuttal exchanges also ran longer in the feedback arm. The abstract's "more than 20,000" refers to randomly selected reviews. The statistics section reports feedback posted to 18,946 of 44,831 reviews (42.3%), with 3,521 selected reviews receiving none, which implies about 22,467 selected in total. That last figure is derived from the excerpt's numbers, not stated in it.
The mechanism is deliberately narrow. The agent fired once per selected review, on first submission, and posted its comments an hour later so reviewers could fix typos first. It looked for three problems: vague or generic critiques, questions that overlooked parts of the paper could answer, and unprofessional statements. Feedback went only to the reviewer and the program chairs, was not shared with authors or area chairs, and was not a factor in decisions. Reviewers were told it came from an LLM, and the system made no direct edits. Reliability tests acted as the gate: feedback was posted only if it passed every test, which kept 829 selected reviews from receiving feedback. The authors treat voluntariness as a design principle: reviewers "could opt out by ignoring the feedback." Each review took about a minute and cost about 50 cents.
Set against the nearest notes, the study asks a different question from the ICML 2026 experiment in Does banning LLM use in peer review change review outcomes?. That study measured whether reviewers obeyed rules on LLM use and found substantial noncompliance under both policies. Here no rule was involved. Uptake was the reviewer's choice, so the 27% figure measures voluntary adoption, not compliance. The agent also sits at a different point in the review chain from Can inference scaling help reviewers catch errors humans miss?. That agent audits manuscripts for flaws human reviewers missed, while this one audits reviews, so its quality measure is informativeness and specificity, not detected errors. Its gated, multi-stage design echoes Can structured pipelines make LLM novelty assessment reliable?, but it differs from Can automated review loops handle AI-generated research at scale?. aiXiv iterates review and refinement in a closed loop, whereas the ICLR agent acted only on the initial review, with no later exchange.
The excerpt does not establish that the revisions were better in any way that matters to authors. "More informative" is a blinded rating whose protocol, sample and effect sizes the excerpt does not give. The 80-word increase is conditional on updating, so it describes the revising minority, not the whole feedback arm. The discussion says feedback-arm reviewers were more likely to change their scores after rebuttal, but gives no figures, and the "significantly longer" wording comes without its test. Nothing here shows changed decisions, and the authors built the system and report its outcomes. The defensible reading is narrower than the headline: a gated, optional feedback layer at scale moved a meaningful minority of reviewers to revise, and blinded raters found the revisions clearer. Whether that improves what authors receive needs an evaluation this excerpt does not contain.
Inquiring lines that read this note 60
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can AI systems perform peer review as effectively as humans?- What specific errors did participants report finding in the AI-generated reviews?
- Did adding AI reviews actually change peer review decisions or paper outcomes?
- Do AI reviews depend more on writing style than scientific merit?
- Should AI research papers require dedicated automated review systems instead of existing journals?
- Does review length bias affect acceptance decisions at major conferences?
- How much does reviewer consistency vary across different papers at NeurIPS?
- Can LLM reviewers catch technical issues that human reviewers miss?
- Can automated review systems catch deep methodological flaws or only surface issues?
- What specific tasks do reviewers use AI for most often?
- Could AI improve peer review rigor and catch human-missed errors?
- How did researchers measure whether GPT-4 and human reviewers identified the same issues?
- Does matching reviewer points actually mean the feedback is accurate or correct?
- Why did rejected papers show higher overlap between GPT-4 and human reviewers?
- Could AI feedback work as a substitute for human peer review entirely?
- Can computational inference scaling catch flaws that human expert reviewers miss?
- Can institutional statements alone correct misconceptions from unreviewed papers?
- Do AI-generated research reviews score papers higher than human reviewers do?
- How often do researchers suspect peer reviews are written by AI?
- Does AI content in reviews correlate with differences in paper quality control?
- Can human reviewers reliably detect AI-written peer review text by sight?
- Can machine review catch flaws in AI-generated work that humans miss?
- How do automated reviewers detect flaws that human experts miss in manuscripts?
- Should AI-generated papers use specialized review venues instead of traditional journals?
- Do peer reviewers actually follow restrictions on using AI tools themselves?
- Are refereed venues also overwhelmed by AI-generated low-quality submissions?
- Can automated systems scale peer review faster than human moderators?
- Can AI reviewers detect deep theoretical flaws that human experts miss?
- Does rhetorical presentation bias reviewers against substantive scientific contributions?
- Does rhetorical quality in reviews influence paper acceptance scores more than content?
- Can agentic AI systems catch flaws in manuscripts that human reviewers consistently miss?
- Why do individual peer reviewers show such low agreement on research merit?
- Could hidden prompts be inserted during review and removed before publication?
- Could automated review systems handle AI-generated research at scale?
- Which feedback loops in AI-mediated review remain unmeasured or rarely observed directly?
- Do shortened peer review timelines correlate with lower quality publications?
- Does an automated reviewer's output actually match human review accuracy?
- Can feeding review scores back into idea generation improve research quality?
- Can automated reviewers actually handle the review load AI creates?
- What role should humans play in reviewing and approving AI-generated research?
- What would a practical reviewer checklist for autonomous research systems need to include?
- Can LLM-generated reference reviews detect machine-written peer review submissions?
- What quality differences exist between flagged and unflagged peer reviews?
- Why do readers rate LLM-edited text more favorably?
- Do humans and LLMs agree on novelty assessment in research?
- Can reviewer-author matching by LLM use amplify biases in acceptance decisions?
- Do reviewer reward badges risk encouraging lenient or superficial reviews?
- Do stricter AI policies actually change how reviewers score manuscripts?
- Do peer review policies banning LLM use actually change reviewer behavior and decisions?
- What happens when conferences enforce bans or limits on reviewer LLM use?
- Can rules against undisclosed LLM use change reviewer behavior without enforcement?
- Do reviewer rules about LLM use in peer review actually get followed?
- Do conference policies banning LLM use actually reduce AI involvement in reviews?
- Why do peer review policies often fail to change actual review scores?
- Do peer reviewers actually follow policies that ban or limit their LLM use?
- How effective are journal policies restricting LLM use in peer review?
- Can peer review policies actually prevent LLM use when compliance is hard to monitor?
- Do metareviewers and regular reviewers use LLMs differently in peer review?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does banning LLM use in peer review change review outcomes?
Can policies restricting or allowing AI tools shift how reviewers score papers and make decisions? This matters because review quality and fairness depend on consistent standards.
contrast: that trial measured compliance with LLM rules; this one measures voluntary uptake of feedback
-
Can inference scaling help reviewers catch errors humans miss?
Explores whether spending extra compute at review time—checking proofs and experiments line by line—can surface deep flaws that evade human expert reviewers, and how this scales with AI-assisted submissions.
contrast: that agent audits manuscripts for flaws; this one audits the reviews themselves
-
Can automated review loops handle AI-generated research at scale?
As AI agents produce papers faster than humans can evaluate them, can a closed-loop automated review system with retrieval-augmented feedback actually improve quality and catch problems traditional peer review misses?
contrast: aiXiv iterates automatically; the ICLR agent gives one feedback pass on the initial review
-
Can structured pipelines make LLM novelty assessment reliable?
Explores whether breaking novelty assessment into extraction, retrieval, and comparison stages helps LLMs align with human peer reviewers and produce more rigorous, evidence-based evaluations.
shared pattern: a multi-stage LLM pipeline, here with reliability tests as a posting gate
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- Stop Automating Peer Review Without Rigorous Evaluation
- LLM-REVal: Can We Trust LLM Reviewers Yet?
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- Can large language models provide useful feedback on research papers? A large-scale empirical analysis
- Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards
- Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?
Original note title
feedback on peer reviews at ICLR 2025 led 27 percent of reviewers to revise and made revisions more informative — a randomized trial by its developers