Do authorship labels change how AI judges evaluate rule violations?
When AI evaluators see a constraint-breaking text, does knowing whether a human or AI wrote it shift their judgment? This tests whether AI judges apply consistent standards or defer to human authority.
The sharpest case in the excerpt is the Lipogram exercise, where Queneau's constraint is to ban the letter "e" entirely. Wright's translation keeps the rule through circumlocution; the GPT-4 version contains 9 instances of the forbidden letter. Here the attribution label moves the two evaluator groups in opposite directions. AI models chose the violating version 28.9% of the time (11 of 38 model runs) when it was labeled AI-generated and 63.9% (23 of 36) when it was labeled human-written, a +35.0 point shift. Human participants chose it 38.7% of the time (12 of 31) when labeled AI-generated and 18.2% (6 of 33) when labeled human-written, a −20.5 point shift.
The excerpt's reading is that AI models "relaxed standards when they believed humans were responsible for constraint violations," while humans stayed "anchored to objective rule compliance." The AI rationales show the leniency in the models' own words. One evaluator said the violating version "successfully captures the essence of the 'Lipogram' style," and another praised it for using "the lipogram more subtly, making it a more authentic representation." The excerpt's summary also says humans "consistently selected Queneau's constraint-adherent version regardless of attribution," but its own human figures show the violating version chosen less when labeled human-written (18.2% against 38.7%). The excerpt does not reconcile the two. The Limitations section adds that the rationales are produced after the choice, so they are the model's account of its selection, not direct evidence of what drove it.
This case extends the rhetorical-sensitivity finding in How much does rhetorical style shift AI review scores?. That note shows reviewer scores moving when only framing changes. The Lipogram is harder to dismiss as a matter of taste, because the violation can be verified by counting letters, and the AI judges still forgave it once a human label was attached. The same authority that Does polished AI output trick audiences into trusting it? attributes to polish appears here to attach to the human label as well; the Discussion describes the result as a violation "reframed as creative latitude." The sibling note, Do authorship labels bias how we judge literary quality?, gives the aggregate pattern that this single exercise illustrates.
What the excerpt does not establish: the Lipogram result rests on 31 and 33 human participants and 38 and 36 AI model runs, and the excerpt shows no other constraint exercise with the same comparison. The Discussion mentions 5,846 paired cases and a coding of "criterion inversions" across the full corpus, but the prevalence results are in the supplementary material and not in this excerpt. Nor does the excerpt show whether the AI leniency reflects learned norms about human creativity or a broader generosity toward plausible authors; the paper's explanation is offered, not tested. The implication is narrow. A verifiable check does not by itself protect an AI judge from authorship labels, and the excerpt supports that claim for one exercise, not as a measured rate across constraints.
Inquiring lines that read this note 15
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How reliably can humans and AI detectors identify machine-generated text?- How much does the human-authorship halo affect AI evaluation across different task domains?
- Can verifiable rule violations protect AI judgment from authorship label bias?
- How often do human annotators mistake human writing for AI-generated text?
- Can AI text detection improve enough to help evaluators make better decisions?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
How much does rhetorical style shift AI review scores?
When manuscripts are rewritten to improve rhetoric while keeping scientific content identical, do LLM reviewers change their scores? Understanding this matters for ensuring AI-assisted peer review evaluates substance, not polish.
same pattern of content-fixed judgment shifts, here on a rule that can be checked by counting letters.
-
Do authorship labels bias how we judge literary quality?
When readers and AI systems know whether text is human- or AI-written, does that label shift their judgments of the same passage? This matters because it tests whether evaluation is based on actual content or on authorship cues.
the sibling note; this Lipogram case is one exercise within the aggregate attribution bias.
-
Does polished AI output trick audiences into trusting it?
When AI generates professional-looking graphs, diagrams, and presentations, do audiences mistake visual polish for analytical depth? This matters because appearance might substitute for actual expertise.
the presentation authority that note describes also attaches to a human label, turning a flaw into "creative latitude."
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- The human-authorship halo: attribution bias in literary style evaluation by humans and AI
- People Overtrust AI-Generated Medical Advice despite Low Accuracy
- Penalizing Transparency? How AI Disclosure and Author Demographics Shape Human and AI Judgments About Writing
- Stop Automating Peer Review Without Rigorous Evaluation
- Humans or LLMs as the Judge? A Study on Judgement Biases
- Do LLMs produce texts with "human-like" lexical diversity?
- GenAI as a Power Persuader: How Professionals Get Persuasion Bombed When They Attempt to Validate LLMs
- When AIs Judge AIs: The Rise of Agent-as-a-Judge Evaluation for LLMs
Original note title
AI evaluators forgave a broken lipogram when they believed a human wrote it, while human judges moved the other way