INQUIRING LINE

If you rewrite a paper's framing but leave its science alone, do AI reviewers' scores still shift?

Does rhetorical quality in reviews influence paper acceptance scores more than content?

This explores whether the way a paper is written (its framing, polish and confident presentation) moves reviewer scores more than the actual science it reports, for both human and AI reviewers.


This explores whether how a paper is written, rather than what it shows, drives the scores reviewers give it. The corpus has no head-to-head study measuring rhetoric against content for human peer reviewers. It does show that presentation moves scores more than most people would expect. The clearest evidence comes from AI reviewers. When researchers rewrote manuscripts' rhetoric but left the scientific content untouched, LLM reviewer scores shifted measurably How much does rhetorical style shift AI review scores?. The biggest swings came from how evidence was framed and how boldly novelty was claimed. Those are exactly the moves authors already make when they 'sell' a paper.

The weakness also extends past phrasing. LLM judges can be fooled by fake references and rich formatting, without any access to the model Can LLM judges be fooled by fake credentials and formatting?. Human readers share a similar habit: they prefer answers with more citations even when the citations are irrelevant Do users trust citations more when there are simply more of them?. Polish also misleads people. Evaluators have rated AI-generated documents as both human-written and better than real human submissions Does polished writing actually signal better quality work?. AI-written documents in particular make readers see more confidence and quality in the writer Does AI writing assistance change how readers perceive the writer?. In practice, polish was enough for one fully AI-generated paper to score 6.33 at an ICLR workshop, meeting the acceptance threshold. Its own authors later judged none of their three AI-generated papers good enough for the main conference track Can AI-generated papers pass peer review undetected?.

In the counterevidence, content shows up as the bigger force. Across more than 125,000 reviews, it looked as if LLM-assisted reviewers favored LLM-written papers. That pattern disappeared once paper quality was controlled for, because LLM papers were mostly the weaker ones Do LLM reviewers actually favor LLM-written papers?. A study from outside peer review makes a similar point: in debates, the voters' prior beliefs predicted who won better than anything about the language did Does what readers believe matter more than what debaters say?. Many 'style effects' turn out to be effects of who is reading or of quality differences that weren't measured. Polish also doesn't always help. In one admissions program, applicants whose essays were improved by AI were admitted at lower rates Does AI essay use hurt admissions chances despite quality gains?.

If the question is about how reviews themselves are written, the evidence points toward specificity mattering more than eloquence. At ICLR 2025, LLM feedback led 27 percent of reviewers to revise their reviews, and blinded raters judged the revised reviews more informative Can LLM feedback help peer reviewers improve their own reviews?. But when ICML 2026 compared banning LLM use in reviewing with allowing limited use, scores and decisions barely changed Does banning LLM use in peer review change review outcomes?. Together these suggest that changing how reviews are written may have little effect on which papers get in.

So here's what the corpus supports. Rhetoric clearly moves scores, especially for AI reviewers, and the effects are big enough to exploit. But in the large real-world datasets, quality still sorts papers once you control for it. The open risk is the one the AI-review study points to: as more reviewing is automated, evaluation becomes the stage where framing gets the most leverage.


Sources 11 notes

How much does rhetorical style shift AI review scores?

Rewriting manuscripts' rhetoric while preserving scientific content moves LLM reviewer scores measurably. Evidence framing and novelty stance produce the largest contrasts; effects vary by reviewer model and interact with the paper's original quality tier.

Can LLM judges be fooled by fake credentials and formatting?

Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.

Do users trust citations more when there are simply more of them?

Analysis of 24,000 Search Arena interactions shows irrelevant citations boost user preference (β=0.273) nearly as much as relevant citations (β=0.285), indicating citation count functions as a decoupled trust heuristic.

Does polished writing actually signal better quality work?

Studies show evaluators perceived AI-generated documents as both human-written and better quality than human submissions. This suggests rhetorical polish misleads judgment and should not serve as a quality signal in evaluation.

Does AI writing assistance change how readers perceive the writer?

A study of 2,939 writers and 11,091 readers found AI assistance shifted every tested dimension—29 total—toward extremism, confidence, quality, agreeableness, and perceived privilege. Distortions were statistically significant and directional, not random noise.

Show all 11 sources
Can AI-generated papers pass peer review undetected?

Sakana AI's end-to-end system produced a paper that scored 6.33 in double-blind ICLR 2025 workshop review, meeting acceptance thresholds, but was withdrawn under pre-agreed protocol. Authors later identified a citation error and judged none of three submissions suitable for main-track publication.

Do LLM reviewers actually favor LLM-written papers?

Across 125,000+ reviews, the apparent favoritism of LLM-assisted reviewers toward LLM papers disappears once paper quality is held constant. LLM papers cluster among weaker submissions, creating a spurious interaction driven by LLM reviewers' general leniency toward lower-quality work.

Does what readers believe matter more than what debaters say?

Analysis of debate corpora shows that political and religious ideology labels of voters outpredict linguistic features when modeling debate outcomes. Language effects observed without reader controls are confounded by audience composition correlated with debate topics.

Does AI essay use hurt admissions chances despite quality gains?

Among 7,500 applications to a public policy master's program, majority of 2025 applicants submitted AI-generated essays despite explicit prohibition. These applicants were admitted at lower rates than similar applicants without detected AI use, despite AI improving essay quality.

Can LLM feedback help peer reviewers improve their own reviews?

A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.

Does banning LLM use in peer review change review outcomes?

A randomized experiment at ICML 2026 found that prohibiting LLM use versus allowing limited use barely changed paper scores, decisions, or reviewer confidence. Meanwhile, substantial fractions of reviewers broke whichever rule they were given.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.