Do authorship labels bias how we judge literary quality?
When readers and AI systems know whether text is human- or AI-written, does that label shift their judgments of the same passage? This matters because it tests whether evaluation is based on actual content or on authorship cues.
Two controlled studies built on Queneau's Exercises in Style (1947) find what the paper calls "systematic pro-human attribution bias." Study 1 gave 556 human participants and 13 AI models literary passages from Queneau and GPT-4-generated versions, under blind, accurately labeled and counterfactually labeled conditions. Humans showed +13.7 percentage points of bias (Cohen's h = 0.28, 95% CI: 0.19–0.37). AI models showed +34.3 points (h = 0.70, 95% CI: 0.62–0.78), which the paper calls a 2.5-fold stronger effect (P<0.001). Study 2 used a 14×14 matrix of evaluators and creators and found the bias across AI architectures at +25.8 points (95% CI: 24.1–27.6%). The paper's summary is that AI systems "systematically devalue creative content when labeled as 'AI-generated' regardless of which AI created it."
The design holds the story content fixed and varies only the attribution label, so the difference is attributed to the label. In the Discussion, the same model judging the same two passages under correct and reversed labels gave opposing assessments: a dialect feature praised as "authentic" under a human label became "exaggerated" under an AI label. The paper reads this as evaluators who "bring learned biases about creative agency to their assessments." The abstract adds that the bias may come from training, "including through the preference signals on which they are aligned." That step is offered as a suggestion; the excerpt does not test the alignment data. The Limitations section also notes that AI explanations are generated after each choice, so they show what a model says about its selection, not the process that produced it.
This sits against the nearest notes in a specific way. How much does rhetorical style shift AI review scores? finds that LLM reviewer scores move when only the rhetoric of a manuscript changes. This excerpt finds a similar movement when only the label changes, which suggests the judges respond to cues about a text more than to what the text reports. Does polished AI output trick audiences into trusting it? argues that polished output borrows authority from its presentation; the excerpt shows that the same text loses credit when labeled machine-written, so the authority in these studies runs on provenance as well as polish. Can humans detect AI text if machines can measure it? is about detection. The excerpt does not test detection; its claim concerns judgments that move once provenance is known.
Several things are not established here. The counterfactual condition is named, but its results are not in the excerpt. Human recruitment and sampling are not described beyond N=556. Study 1 uses one generator (GPT-4), one source narrative and minimal prompting, and the human reference set is one 1947 text in one English translation, across thirty selected exercises. The effect sizes therefore describe this paradigm, a style task on one retold story, and are narrower than a general law of how AI judges AI work. What the excerpt supports is that label-sensitive judgment is a measurable risk for creative evaluation, and that the AI shift is larger than the human one in this setup.
Inquiring lines that read this note 22
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How reliably can humans and AI detectors identify machine-generated text?- Can verifiable rule violations protect AI judgment from authorship label bias?
- Can a classifier distinguish machine-written text from poor human writing?
- What false-positive rate would indicate the classifier harms legitimate human writers?
- Do human readers still recognize authors after heavy AI rewriting?
- Does AI assistance distort how readers perceive writer identity and demographics?
- Does polish in writing borrow authority that only expertise should carry?
- Can readers actually distinguish AI text from human writing?
- Why do people rate AI-written text as better than human writing?
- Can readers reliably distinguish AI-written abstracts from human-written ones?
- Does AI-written text score higher because of presentation alone or judgment shift?
- Do human reviewers detect rhetorical polish as a sign of AI authorship?
- Can literary quality be measured precisely enough to train models?
- How does salience of AI involvement shape judgments at the moment of reading?
- How much does knowing about AI use actually change how readers judge text?
- Can writers claim authorship without feeling cognitive ownership of the work?
- Do writers claim authorship without feeling they wrote the words?
- What aspects of authenticity matter most to readers versus writers?
- Do writers experience felt authorship differently from authorship they claim?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
How much does rhetorical style shift AI review scores?
When manuscripts are rewritten to improve rhetoric while keeping scientific content identical, do LLM reviewers change their scores? Understanding this matters for ensuring AI-assisted peer review evaluates substance, not polish.
both find LLM judges moving on cues about a text with its content fixed; this excerpt varies the label, not the rhetoric.
-
Does polished AI output trick audiences into trusting it?
When AI generates professional-looking graphs, diagrams, and presentations, do audiences mistake visual polish for analytical depth? This matters because appearance might substitute for actual expertise.
extends the authority argument: the same text loses credit when labeled machine-written, so provenance also carries weight.
-
Can humans detect AI text if machines can measure it?
AI-generated text shows measurable differences from human writing across multiple linguistic dimensions, yet human judges consistently fail to identify it. Why does the gap between what is measurable and what is perceptible exist?
the excerpt does not test detection; it shows judgments moving once provenance is disclosed, a separate question from perceptibility.
-
Do authorship labels change how AI judges evaluate rule violations?
When AI evaluators see a constraint-breaking text, does knowing whether a human or AI wrote it shift their judgment? This tests whether AI judges apply consistent standards or defer to human authority.
Qualifies the human-halo: on a rule-breaking lipogram, human judges reversed direction while AI evaluators forgave the violation when told a human wrote it
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- The human-authorship halo: attribution bias in literary style evaluation by humans and AI
- Understanding Reader Perception Shifts upon Disclosure of AI Authorship
- Penalizing Transparency? How AI Disclosure and Author Demographics Shape Human and AI Judgments About Writing
- The AI Ghostwriter Effect: When Users Do Not Perceive Ownership of AI-Generated Text But Self-Declare as Authors
- Do LLMs produce texts with "human-like" lexical diversity?
- Measuring and Mitigating Persona Distortions from AI Writing Assistance
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
- Beyond "Made with AI": Visualizing Provenance Density to Mitigate the Transparency Penalty
Original note title
authorship labels tilt literary style judgments toward human authors, and AI evaluators show a 2.5-fold stronger tilt — the human-authorship halo