Can we judge text quality without knowing who wrote it?
Does the concept of 'slop' work as a quality judgment independent of whether a machine or human authored the text? This matters because current AI detection often conflates two separate questions: origin and quality.
The paper separates two questions that detection work tends to fuse: whether a machine wrote a text, and whether the text reads as slop. It says slop "differs from AI-text detection in general, and can be applied to any text source (whether AI-written or not)." It adds that "Human writing can also read as 'slop'", although it adopts the definition and focuses on "(seemingly) LLM-generated texts." Its stated aim is to characterize "qualities of texts that contribute to them being categorized as 'slop,'" which it suggests may explain cases where humans mistake human-written text for AI-generated text.
The reasoning turns on what the judgment is about. Detectors such as DetectGPT and Binoculars score the likelihood of AI origin and, as the paper cites them, "report high discriminant performance (0.95 AUROC)." The paper states that "our taxonomy and annotations diverge from those used for AI-text detection in general." Its target is still framed as "stylistic patterns unique to LLM writing," but the definition it works from is built on observable quality dimensions that any text can have. The binary judgments that anchor the work are made by annotators and, per the abstract, "correlate with latent dimensions such as coherence and relevance."
This sharpens Can human judges detect measurable differences in AI text?, which finds that LLM text differs measurably in lexical diversity and that human judges cannot identify it. Slop judgments are a different kind of human response: people make them, and they are framed around quality rather than origin. That is also the contrast with Can humans detect AI text if machines can measure it?. The two notes are about different targets of the same kind of measurement, origin in one case and quality in the other. The excerpt does not report whether its annotators could tell which texts were AI-written, so the comparison stops at the framing.
The excerpt does not establish how slop judgments relate to origin in practice. It reports no test of the framework against detectors, and it gives no figures on how often human-written text is judged slop. Its own target, LLM-specific style, sits in some tension with a definition that applies to any text, and the excerpt does not resolve that tension. The implication is that the category can be used to critique writing wherever it comes from, but a claim that a text judged slop was machine-written goes beyond what the excerpt supports.
Inquiring lines that read this note 6
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do interpretive frames override surface features in text comprehension? How reliably can humans and AI detectors identify machine-generated text? Can readers reliably distinguish AI-written text from human writing? Why does polished AI output gain credibility despite fundamental verifiability problems?Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can human judges detect measurable differences in AI text?
Research shows LLM text differs statistically across six lexical dimensions, but human readers—even experts—cannot reliably identify which texts are AI-generated. Why does measurement succeed where human perception fails?
origin measurement in that note; this paper asks whether a text reads as slop, which applies to human text too.
-
Can humans detect AI text if machines can measure it?
AI-generated text shows measurable differences from human writing across multiple linguistic dimensions, yet human judges consistently fail to identify it. Why does the gap between what is measurable and what is perceptible exist?
contrast: that note concerns detecting origin, while slop judgments are made by people and framed around quality.
-
What dimensions make text feel like AI slop?
Can we break down the vague notion of AI slop into measurable components? Researchers coded expert definitions to find which specific text properties people associate with low-quality generated writing.
sibling: the axes that structure the quality judgment this note describes.
-
Does polished writing actually signal better quality work?
When evaluators judge applications and manuscripts, does rhetorical sophistication predict merit, or does it distract from verifiable evidence of competence and rigor?
qualifies: B finds evaluators rated AI text as human and better, so polish-based quality readings may depend on perceived authorship
-
Do authorship labels bias how we judge literary quality?
When readers and AI systems know whether text is human- or AI-written, does that label shift their judgments of the same passage? This matters because it tests whether evaluation is based on actual content or on authorship cues.
qualifies: authorship labels tilt literary style judgments, so quality readings of text may depend on authorship labels
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Measuring AI "Slop" in Text
- Penalizing Transparency? How AI Disclosure and Author Demographics Shape Human and AI Judgments About Writing
- Stop Automating Peer Review Without Rigorous Evaluation
- Do LLMs produce texts with "human-like" lexical diversity?
- "That's AI Slop, You Bot!" Studying Accusations, Evidence, and Credibility in Online Discourse Towards LLM-Generated Comments
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
- Pangram Predicts 21% of ICLR Reviews are AI-Generated
Original note title
slop is a quality judgment that can apply to any text, not a test of machine authorship like AI-text detection