INQUIRING LINE

When people call something 'AI slop,' are they spotting machine text, or just saying it doesn't belong?

How do slop judgments correlate with actual AI detection performance in practice?

This explores whether people's 'that's AI slop' calls track which text is actually machine-written, and what the corpus says about where real detection signals live.


This explores whether calling something 'AI slop' is a reliable way to spot machine-written text, or whether it measures something else. The most direct answer in the collection is no. A matched-control study of 25 million Hacker News and Reddit comments found that the prose features that actually separate AI text from human text do not predict which comments get accused of being slop Do AI slop accusations actually detect AI text?. The accusation behaves more like a social signal, a way of saying 'this doesn't belong here,' than like a screening tool. Calling something slop says more about community norms than about where the text came from.

This fits a broader finding about human judgment. A review of 30 studies found that people's ability to tell AI content from human content in text, images and voice sits at about chance, and it hasn't improved as generators have gotten more realistic Can people reliably spot content made by AI?. So slop accusations are not a weak version of a skill people have. There's little underlying skill for them to approximate. People react to vibes like blandness, hedging or a certain polish, and those vibes don't reliably mark machine authorship.

The less obvious point is where the real signal seems to be. Research on AI-written fiction found that surface style is the wrong place to look. A detector using only story-level choices, such as how much agency characters have and how the timeline is structured, separated AI from human stories with 93% accuracy. It kept almost all of that accuracy after stylistic cues were removed Can AI stories be detected without analyzing writing style?. Those structural habits are hard to disguise because fixing them means rewriting the piece, not tweaking word choice. That is roughly the reverse of how people judge slop: humans look at phrasing, while the dependable differences are in the larger decisions a reader rarely notices consciously.

Automated judges have a similar weakness. LLM evaluators reliably give higher scores to responses with fake citations or rich formatting, whatever the content quality Can LLM judges be tricked without accessing their internals?. Reward models have the same problem. Splitting quality into specific, checkable criteria helps them stop overfitting to superficial artifacts Can breaking down instructions into checklists improve AI reward signals?. Human or machine, any judge working from overall impressions tends to lock onto surface features. Explicit, structural criteria are what improve reliability.

The gap: this collection has no study that directly measures how well individual people's slop calls match ground-truth labels, or that compares them with automated detectors on the same text. The case above is assembled from adjacent evidence. That evidence is consistent, but it doesn't settle the question.


Sources 5 notes

Do AI slop accusations actually detect AI text?

A matched-control study of 25 million Hacker News and Reddit comments found that prose features distinguishing AI from human text do not predict which comments get accused as slop. The label functions as social regulation rather than accurate screening.

Can people reliably spot content made by AI?

A 30-study systematic review found that humans cannot reliably distinguish AI-generated from human-created content across text, image, and voice modalities. Accuracy generally clusters around chance and has not kept pace with improvements in AI realism.

Can AI stories be detected without analyzing writing style?

StoryScope achieved 93.2% accuracy separating AI from human fiction using only discourse-level features like character agency and chronological structure, retaining 97% of performance while eliminating stylistic cues. These structural choices resist humanization because they require rewrites, not surface edits.

Can LLM judges be tricked without accessing their internals?

Research shows LLM evaluators systematically score higher when responses include fake references or rich formatting, independent of content quality. These biases are exploitable without model access, undermining AI benchmark credibility.

Can breaking down instructions into checklists improve AI reward signals?

RLCF and RaR methods decompose instruction quality into verifiable sub-criteria, improving performance on benchmarks like FollowBench and HealthBench. This decomposition principle reduces overfitting to superficial artifacts that plague holistic reward models.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.