Can you really tell when writing was made by AI, or do you only feel sure you can?
What makes readers suspect AI involvement in academic writing they evaluate?
This explores what signals lead readers, reviewers or evaluators to suspect that a piece of writing they're judging was produced with AI. The corpus has almost no material on academic papers specifically, so this answer draws on nearby settings: admissions essays, online comments, fiction and news.
This explores what signals lead readers, reviewers or evaluators to suspect that a piece of writing they're judging was produced with AI. The corpus doesn't have research on reviewers of academic papers. What it does have is evidence from neighboring settings, and it points somewhere unexpected: the cues readers rely on are often weaker than they think, and AI writing leaves its clearest traces in places readers don't usually look.
Start with the setting closest to academic evaluation. Admissions officers in one experiment could often tell AI-written essays from human ones, and they rated essays they believed were AI-generated lower Do admissions officers penalize essays they suspect are AI-written?. But elsewhere, the confidence behind an accusation doesn't match the evidence. When researchers looked at online comments that people had accused of being AI-written, those comments lacked the features that actually separate AI text from human text. That suggests the accusations worked more as gatekeeping than detection, and the people harmed were the real human writers who weren't believed Do unfounded AI accusations harm human writers instead?. In another study, readers simply couldn't tell AI-assisted writing from writing done alone Do readers value writing authenticity they cannot detect?.
So what does AI writing actually change? One large study (2,939 writers, 11,091 readers) found that AI assistance shifted how readers saw the writer on all 29 traits measured. The writers came across as more confident, more extreme, more agreeable, and higher in quality Does AI writing assistance change how readers perceive the writer?. They also came across as more privileged: more educated, higher-income, and more likely to be native English speakers than they really were. Researchers call this 'identity laundering': a distinctive personal voice gets flattened into a generic, polished one Does AI writing make authors seem more privileged than they are?. And these shifts reach readers almost untouched, because writers edited AI-generated paragraphs only 23% of the time, and even then barely changed them Do writers actually edit AI-generated text before publishing?. None of these studies tested whether that polished, self-assured voice is what makes readers suspicious. But it's a reasonable guess about the 'feel' that evaluators react to. It also suggests a risk: a writer whose natural voice is unusually fluent or formal could be suspected for the same reasons.
The strongest detection evidence comes from structure, not wording. In fiction, a system called StoryScope told AI stories from human ones with 93% accuracy using only narrative choices like how characters act and how events are ordered. It kept nearly all of that accuracy after every stylistic cue was removed Can AI stories be detected without analyzing writing style?. For academic writing, the parallel would be how an argument is built, not which words appear in it. That's a guess, though: the corpus hasn't tested it on papers. Whether heavy rewriting also fools automated detectors is still an open question Do rewrites that hide authorship also fool AI detectors?.
Finally, suspicion matters partly because of what happens after it. When readers are told AI was used, trust drops, most sharply in personal writing How does revealing AI authorship change reader trust?. The penalty on a straightforward news article was small but consistent, and it showed up even when the raters were LLMs Does disclosing AI assistance make readers trust articles less?. Readers with more AI literacy penalize less Does AI literacy reduce the damage from AI disclosure?. And readers think disclosure is more necessary than writers do, especially when AI text is pasted in directly and couldn't easily be replaced Do readers and writers differ on AI disclosure necessity?. The takeaway: what readers suspect, and how harshly they judge it, depends as much on their expectations as on anything in the text.
Sources 12 notes
In an experiment, admissions officers could often discriminate AI from human essays and rated essays they believed to be AI-generated lower than those believed human-written. The authors frame this as a plausible explanation for the observed admissions penalty, though the link remains proposed rather than directly measured.
Accused comments lack features that distinguish AI text from human writing, suggesting accusations function as gatekeeping rather than detection. This inverts the AI-as-perpetrator framing, placing harm at the receiving side through reader skepticism.
Hwang et al. found that readers could not distinguish AI-assisted from solo-written work and showed positive attitudes toward AI use. However, the study did not test whether readers would value process authenticity if disclosure occurred or if they could perceive it.
A study of 2,939 writers and 11,091 readers found AI assistance shifted every tested dimension—29 total—toward extremism, confidence, quality, agreeableness, and perceived privilege. Distortions were statistically significant and directional, not random noise.
Writers using AI assistance were perceived as significantly more educated (5.3×), higher-income (4.4×), native English speakers (4.1×), and white (1.1×). This demographic distortion compresses distinctive voice markers into a generic privileged persona, creating what researchers call identity laundering.
Show all 12 sources
Writers edited AI-generated paragraphs only 23% of the time, with edits averaging 96% similarity to the original. This means AI's opinionated and distorted voice propagates with minimal human filtering before publication.
StoryScope achieved 93.2% accuracy separating AI from human fiction using only discourse-level features like character agency and chronological structure, retaining 97% of performance while eliminating stylistic cues. These structural choices resist humanization because they require rewrites, not surface edits.
The paper asserts that rewritten messages evade AI-text detectors but provides no detector experiments, only attribution results showing stylistic convergence. The double erasure claim needs direct empirical testing.
A study of 261 readers found that disclosing AI authorship consistently lowered perceived trustworthiness, caring, and likability, with the steepest drops in interpersonal writing like personal interaction. Readers saw AI as incapable of genuine empathy, viewing its use as a violation of social expectations.
Both human raters (n=1,970) and LLM raters (n=2,520) scored an identical news article lower when it included an AI disclosure statement, but the penalty was small—less than 0.15 points on a 7-point scale.
In a 261-person study, readers with higher self-reported AI literacy showed smaller negative shifts in perception after learning AI was used, and some expressed positive attitudes toward AI use. Literacy appears to act as a boundary condition on the broader disclosure penalty.
A 727-person vignette study found readers consistently rated AI disclosure as more necessary than writers did. Disclosure seemed most necessary when AI text was directly incorporated and irreplaceable, while writer effort had no effect on these judgments.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Understanding Reader Perception Shifts upon Disclosure of AI Authorship
- What Influences Readers' and Writers' Perceived Necessity of AI Disclosure?
- Penalizing Transparency? How AI Disclosure and Author Demographics Shape Human and AI Judgments About Writing
- Measuring and Mitigating Persona Distortions from AI Writing Assistance
- The Assistant Erased You: Measuring Loss of Authorship Signals in AI-Mediated Communication
- "It was 80% me, 20% AI": Seeking Authenticity in Co-Writing with Large Language Models
- "That's AI Slop, You Bot!" Studying Accusations, Evidence, and Credibility in Online Discourse Towards LLM-Generated Comments
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries