People can't reliably tell AI-written comments from human ones, yet AI text carries measurable fingerprints that readers miss.
Can readers distinguish machine-generated text from human-written comments?
This explores whether ordinary readers, including experts, can tell AI-written text from human writing in places like comment threads and online posts, and what actually happens when people try.
This explores whether people reading comments, posts and similar short texts can tell which ones a machine wrote. The short answer from the corpus is no, at least not reliably. The more surprising finding is that the text does carry real, measurable signs of machine authorship. Readers just don't notice them, and the people who make accusations aren't reacting to them either.
Start with human judgment. A review of 30 studies found that people's accuracy at spotting AI content in text, images and voice sits around chance, and it hasn't improved as the models have Can people reliably spot content made by AI?. Expertise doesn't help much. ML researchers reading research abstracts couldn't reliably pick out the LLM-written ones and tended to assume a human was involved in all of them Can readers tell LLM abstracts from human ones?. In another study, trained linguists also failed Can human judges detect measurable differences in AI text?.
The differences are there, though. Statistical analysis shows AI text differs from human text on six measures of vocabulary use, such as how varied the word choice is and how evenly words are spread. Newer models drift further from human patterns even as people find them harder to spot Can humans detect AI text if machines can measure it?. Machines can pick up these signals easily. Simple, readable linguistic features identified LLM-written counter-arguments on Reddit's r/ChangeMyView with 99% accuracy. The giveaways included an overly polished, textbook style of argument and a habit of going along with whatever the prompt set up Can simple linguistic features detect AI-written arguments?. In fiction, AI stories can be identified from story structure alone, such as how characters act and how events are ordered, even with all surface style removed. That kind of signal is hard to hide because fixing it means rewriting the story, not polishing the prose Can AI stories be detected without analyzing writing style?.
So what are people reacting to when they call a comment "AI slop"? A matched comparison of 25 million Hacker News and Reddit comments found that the features that really separate AI from human prose don't predict which comments get accused. The label works more like social gatekeeping, a way of policing tone or belonging, than like detection Do AI slop accusations actually detect AI text?. Audiences don't seem to penalize the machine style much either. AI-written Reddit comments, with their recognizably warm, assistant-like tone, got about as much engagement as human comments and sometimes more Does machine-generated text get penalized in online engagement?. On Medium, posts a detector classified as AI-generated got about half the likes of posts it classified as human-written. The authors still described that gap as relatively small Do readers engage less with AI-generated social media posts?.
The part you might not expect: "human-written" is itself becoming a blurry category. Writers who use AI assistance keep its suggestions almost unchanged. They edited AI paragraphs only 23% of the time, and their edits left the text about 96% the same Do writers actually edit AI-generated text before publishing?. That assistance changes how readers see the writer on all 29 traits measured, making the writer seem more extreme, more confident and more agreeable Does AI writing assistance change how readers perceive the writer?. The practical question may be less "was this written by a machine?" and more "how much did a machine shape the person I think I'm reading?" Readers can't answer that one by eye either.
Sources 11 notes
A 30-study systematic review found that humans cannot reliably distinguish AI-generated from human-created content across text, image, and voice modalities. Accuracy generally clusters around chance and has not kept pace with improvements in AI realism.
Readers with ML expertise struggle to identify LLM-generated content reliably, tending to assume human involvement across all abstract types. However, LLM-edited abstracts received highest clarity ratings and were preferred 55% of the time when authorship was disclosed.
Six-dimension MANOVA analysis confirms significant differences between ChatGPT and human writing across vocabulary volume, abundance, variety, evenness, disparity, and dispersion. Despite these robust statistical differences, human judges including linguists and NLP researchers fail to reliably distinguish AI from human text.
LLM-generated text differs significantly on six lexical diversity dimensions, confirmed through statistical analysis across multiple models. Yet human judges, including trained linguists, cannot reliably detect these differences—and newer models diverge further while becoming harder to spot.
General linguistic features combined with argument-quality measures achieved 99% accuracy detecting LLM-generated counter-arguments on r/ChangeMyView, matching heavyweight neural detectors while remaining computationally cheap and transparent. LLMs produce detectable stylistic signatures: accommodation to prompts and textbook-quality argument markers that humans don't replicate.
Show all 11 sources
StoryScope achieved 93.2% accuracy separating AI from human fiction using only discourse-level features like character agency and chronological structure, retaining 97% of performance while eliminating stylistic cues. These structural choices resist humanization because they require rewrites, not surface edits.
A matched-control study of 25 million Hacker News and Reddit comments found that prose features distinguishing AI from human text do not predict which comments get accused as slop. The label functions as social regulation rather than accurate screening.
A Reddit measurement found that machine-generated comments convey assistant-style warmth and status-giving, yet receive engagement levels often indistinguishable from human-authored content and sometimes higher, suggesting the stylistic difference carries no penalty.
AI-labeled posts on Medium averaged 69.15 likes versus 127.59 for human-labeled posts, with similar gaps in comments across all follower groups. The paper calls this gap relatively small and suggests AI content still appeals to users.
Writers edited AI-generated paragraphs only 23% of the time, with edits averaging 96% similarity to the original. This means AI's opinionated and distorted voice propagates with minimal human filtering before publication.
A study of 2,939 writers and 11,091 readers found AI assistance shifted every tested dimension—29 total—toward extremism, confidence, quality, agreeableness, and perceived privilege. Distortions were statistically significant and directional, not random noise.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- The human-authorship halo: attribution bias in literary style evaluation by humans and AI
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
- Is it Cake or is it AI? A Systematic Review of Human Uncertainty in Distinguishing Generative Artificial Intelligence Content
- Measuring AI "Slop" in Text
- Understanding Reader Perception Shifts upon Disclosure of AI Authorship
- Machines in the Crowd? Measuring the Footprint of Machine-Generated Text on Reddit
- Linguistic markers of inherently false AI communication and intentionally false human communication: Evidence from hotel reviews
- Measuring and Mitigating Persona Distortions from AI Writing Assistance