SYNTHESIS NOTE
Topics›Expertise in the Age of AI Content›this note

Can readers tell LLM abstracts from human ones?

Do readers with ML expertise reliably distinguish human-written, LLM-generated, and LLM-edited research abstracts? Understanding this matters for evaluating whether readers can serve as effective gatekeepers against LLM content.

Synthesis note · 2026-10-06 · sourced from Expertise in the Age of AI Content

Akpinar et al. report that readers with machine learning expertise do not reliably separate human-written, LLM-generated and LLM-edited research abstracts. The excerpt says participants "struggle to reliably identify LLM-generated content," and the discussion finds them "tending instead to assume some degree of human involvement" in all three types, with "a baseline suspicion that LLMs were involved across all abstracts." The types were not treated alike, though. With authorship disclosed, LLM-edited abstracts "received the highest clarity ratings (β= 1.383, p< .001) and were selected by 55% of participants when authorship was disclosed, compared to 27-28% for human-written and LLM-generated alternatives."

The authors trace these judgments to heuristics they list as "completeness, clarity, credibility, engagement, and writing conventions," and they conclude that these cues "prove systematically unreliable." Readers valued LLM editing because it "achieved clarity without sacrificing substance"; one participant noted that "LLM introduced linguistic clarity and cohesiveness." LLM-generated abstracts drew criticism for "information overload without focus." The advantage was conditional: LLM-edited versions were "strongly preferred, but only when their LLM authorship level is revealed."

This is a reader-side case of a pattern in How much does rhetorical style shift AI review scores?, where rewriting presentation without changing reported content moved LLM reviewers' scores. The lever here is clarity editing and the readers are human, but both show presentation moving judgments of the same underlying science. The clarity preference also fits the account in Does polished AI output trick audiences into trusting it?, where polish stands in for expertise. The excerpt is consistent with that account but does not test whether readers checked substance. The disclosure half of the finding, which the same survey reports separately, is in Do reader judgments reflect actual authorship or just their beliefs?.

The excerpt does not give the number of participants, how they were recruited, or how often readers identified the true authorship type, so "do not reliably" is the strongest statement it supports; no accuracy rate appears in it. The 55% and 27-28% figures are selection shares under the disclosed condition of one survey, and the excerpt does not say what the β is measured against. Nothing here shows that LLM-edited abstracts are more accurate or more trustworthy than human ones; it shows that readers given these texts rate the edited version as clearer. At this strength, the implication is that a reader's sense of an abstract's quality is a weak check on whether an LLM was involved, so rules that expect readers to catch LLM text in summaries rest on thin ground.

Inquiring lines that read this note 49

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How reliably can humans and AI detectors identify machine-generated text? Does disclosing AI authorship change how audiences evaluate the writing? How can we detect and account for LLM involvement in academic writing? Can readers reliably distinguish AI-written text from human writing? Can AI systems perform peer review as effectively as humans? How do hallucinated citations emerge in AI scholarly output? Do restrictions on reviewer LLM use actually shape peer review behavior? How can AI systems reliably guide voters without introducing political bias? How do clinicians calibrate trust in AI medical recommendations?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
18 direct connections · 112 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

readers could not reliably tell LLM from human research abstracts, yet LLM-edited abstracts were rated most favorably