INQUIRING LINE

Readers can't reliably tell an AI-written research abstract from a human one, and they often prefer the AI-edited version.

Can readers reliably distinguish AI-written abstracts from human-written ones?

This explores whether people, including experts, can tell when a research abstract (or similar text) was written by an AI instead of a person, and what that ability, or lack of it, means in practice.


This explores whether readers can tell AI-written abstracts from human ones, and what that means for trusting what you read. The short answer is no, not reliably. Even readers with machine-learning expertise struggle, and they tend to guess 'a human was involved' no matter who wrote the text Can readers tell LLM abstracts from human ones?. The twist is that abstracts edited by an LLM got the highest clarity ratings and were preferred 55% of the time even when readers were told who wrote them. So the AI version isn't slipping by unnoticed despite being worse. Readers often like it better.

Abstracts aren't a special case. A review of 30 studies found that people's accuracy at spotting AI content clusters around chance across text, images and voice, and it hasn't kept up as AI output has become more realistic Can people reliably spot content made by AI?. The strangest part is that AI text really is different. Statistical analysis finds clear gaps between ChatGPT and human writing on six measures of vocabulary variety. Yet linguists and NLP researchers still can't perceive those gaps, and newer models drift further from human patterns while becoming harder to catch Can humans detect AI text if machines can measure it? Can human judges detect measurable differences in AI text?. The signal is there. Human perception just isn't tuned to it.

Machines do much better when they look in the right place. Simple, transparent linguistic features hit 99% accuracy at flagging LLM-written arguments on Reddit. Two of the giveaways are a habit of mirroring the prompt and 'textbook-quality' argument markers that people rarely use Can simple linguistic features detect AI-written arguments?. In fiction, AI stories can be caught at 93% accuracy from narrative choices alone, such as how characters act and how time is ordered. Those choices hold up against disguise because changing them takes a rewrite, not a polish Can AI stories be detected without analyzing writing style?. A related clue is rhetorical: AI prose is grammatically smooth but avoids taking an evaluative stance, which makes it coherent but argumentatively flat Why does AI writing sound generic despite being grammatically correct?. That flatness may be exactly what makes an abstract read as 'clear'.

In science, the stakes go beyond abstracts. One of three fully AI-generated papers scored above the acceptance threshold at a double-blind ICLR workshop. The authors later found a citation error in it and judged none of the three good enough for the main conference Can AI-generated papers pass peer review undetected?. Meanwhile, the label can matter more than the text itself. Human judges rated identical passages 13.7 points higher when told a human wrote them, and AI evaluators showed a bias 2.5 times stronger Do authorship labels bias how we judge literary quality?. Putting these together, readers can't tell the difference by reading, but they react strongly once they're told.

Here is what you might not have expected to want to know. Detection may be the wrong frame, because AI help changes the writer too. In a study of nearly 3,000 writers, AI assistance shifted how readers saw the author on all 29 measured traits, toward seeming more confident, more extreme and more agreeable Does AI writing assistance change how readers perceive the writer?. Writers edited AI suggestions only 23% of the time, and those edits barely changed the text Do writers actually edit AI-generated text before publishing?. So the more useful question may be less 'was this written by AI?' and more 'whose voice am I actually hearing?'


Sources 11 notes

Can readers tell LLM abstracts from human ones?

Readers with ML expertise struggle to identify LLM-generated content reliably, tending to assume human involvement across all abstract types. However, LLM-edited abstracts received highest clarity ratings and were preferred 55% of the time when authorship was disclosed.

Can people reliably spot content made by AI?

A 30-study systematic review found that humans cannot reliably distinguish AI-generated from human-created content across text, image, and voice modalities. Accuracy generally clusters around chance and has not kept pace with improvements in AI realism.

Can humans detect AI text if machines can measure it?

LLM-generated text differs significantly on six lexical diversity dimensions, confirmed through statistical analysis across multiple models. Yet human judges, including trained linguists, cannot reliably detect these differences—and newer models diverge further while becoming harder to spot.

Can human judges detect measurable differences in AI text?

Six-dimension MANOVA analysis confirms significant differences between ChatGPT and human writing across vocabulary volume, abundance, variety, evenness, disparity, and dispersion. Despite these robust statistical differences, human judges including linguists and NLP researchers fail to reliably distinguish AI from human text.

Can simple linguistic features detect AI-written arguments?

General linguistic features combined with argument-quality measures achieved 99% accuracy detecting LLM-generated counter-arguments on r/ChangeMyView, matching heavyweight neural detectors while remaining computationally cheap and transparent. LLMs produce detectable stylistic signatures: accommodation to prompts and textbook-quality argument markers that humans don't replicate.

Show all 11 sources
Can AI stories be detected without analyzing writing style?

StoryScope achieved 93.2% accuracy separating AI from human fiction using only discourse-level features like character agency and chronological structure, retaining 97% of performance while eliminating stylistic cues. These structural choices resist humanization because they require rewrites, not surface edits.

Why does AI writing sound generic despite being grammatically correct?

AI text uses manner nouns and anaphoric references that are descriptively neutral, while human writers use status and evidential nouns that carry evaluative weight. This produces organizationally coherent but argumentatively inert prose.

Can AI-generated papers pass peer review undetected?

Sakana AI's end-to-end system produced a paper that scored 6.33 in double-blind ICLR 2025 workshop review, meeting acceptance thresholds, but was withdrawn under pre-agreed protocol. Authors later identified a citation error and judged none of three submissions suitable for main-track publication.

Do authorship labels bias how we judge literary quality?

Human judges rated identical passages 13.7 percentage points higher when labeled human-authored; AI models showed a 2.5-fold stronger bias at 34.3 points. The effect persists across AI architectures, suggesting evaluators respond to provenance cues rather than text quality alone.

Does AI writing assistance change how readers perceive the writer?

A study of 2,939 writers and 11,091 readers found AI assistance shifted every tested dimension—29 total—toward extremism, confidence, quality, agreeableness, and perceived privilege. Distortions were statistically significant and directional, not random noise.

Do writers actually edit AI-generated text before publishing?

Writers edited AI-generated paragraphs only 23% of the time, with edits averaging 96% similarity to the original. This means AI's opinionated and distorted voice propagates with minimal human filtering before publication.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.