INQUIRING LINE

When an AI judge is told a human wrote the text, it grades more generously, while human judges get stricter.

How much does the human-authorship halo affect AI evaluation across different task domains?

This explores whether believing a human wrote something makes AI judges score it more generously, and whether that bias changes from one kind of task to another. The corpus answers the first half and has almost nothing on the second.


This explores whether believing a human wrote something makes AI judges score it more generously, and whether that bias changes from one kind of task to another. The corpus can't answer the cross-domain part: it has one direct experiment, in a single domain. That experiment is striking, though. AI evaluators were asked to judge a lipogram, a text written without using a particular letter, that broke its own rule. When told a human wrote it, the AI judges picked it 35 percentage points more often. Human judges did the opposite and became 20 points stricter under the same label Do authorship labels change how AI judges evaluate rule violations?. So the halo isn't just a quirk of AI judges. The AI judges and the human judges moved in opposite directions. The AI relaxed its standards for 'human' work, while the people held more tightly to whether the rule was followed.

The main finding here is that a label alone can sway a judge, even on a task with a clear answer. A lipogram is close to the most checkable creative task there is: either the forbidden letter appears or it doesn't. If an authorship label can move an AI judge there, you'd expect it to matter at least as much on looser tasks like essays, fiction or advice, where nothing so simple can be checked. That's an inference, not something the corpus tests. Still, it points to where to worry most: open-ended domains, where there's little objective fact for the judge to fall back on.

Labels matter so much partly because nobody can check them by eye. A review of 30 studies found that people tell AI-made from human-made text, images and voice at roughly chance level Can people reliably spot content made by AI?. In practice, then, the label is often the only authorship signal a judge has. A related pattern shows up with confidence. Users in every language studied follow confident-sounding AI answers whether or not they're accurate Do users worldwide trust confident AI outputs even when wrong?. Both studies show judges, human or AI, reacting to a signal attached to the content more than to the content itself.

There's a way around the label problem. AI fiction can be identified with 93% accuracy from story-level choices alone, such as how much agency characters have and how the timeline is structured, with no reliance on writing style Can AI stories be detected without analyzing writing style?. That suggests evaluation could rest on features of the work itself rather than on claimed authorship. On the judging side, agent-based evaluators that actively gather evidence cut the drift between different judges' verdicts from 31% to under 1% Can agents evaluate AI outputs more reliably than language models?. That fits the lipogram result: a judge that has to check the evidence is harder to sway with a story about who wrote the text.

What the corpus doesn't have is a comparison across domains like code, math, factual answers and persuasive writing. It also lacks evidence on whether the halo shrinks as tasks become more verifiable, which is the natural next question. One adjacent finding complicates the idea of 'human authorship' itself. People feel more ownership of AI-drafted text the more they steer it, and personalizing the model doesn't change that Does user control over AI text shape feelings of ownership?. As more writing becomes a human–AI blend, a simple human/AI label may say less and less about who actually made the text.


Sources 6 notes

Do authorship labels change how AI judges evaluate rule violations?

AI models chose a rule-breaking lipogram 35 percentage points more often when told a human wrote it, while human judges chose it 20 points less in that condition. The shift suggests AI may relax standards for human work while humans anchor to objective compliance.

Can people reliably spot content made by AI?

A 30-study systematic review found that humans cannot reliably distinguish AI-generated from human-created content across text, image, and voice modalities. Accuracy generally clusters around chance and has not kept pace with improvements in AI realism.

Do users worldwide trust confident AI outputs even when wrong?

Cross-linguistic research shows users in every language trust confident AI outputs even when inaccurate. While confidence expression varies by language, users everywhere track confidence signals rather than accuracy, making overconfident errors systematically followed.

Can AI stories be detected without analyzing writing style?

StoryScope achieved 93.2% accuracy separating AI from human fiction using only discourse-level features like character agency and chronological structure, retaining 97% of performance while eliminating stylistic cues. These structural choices resist humanization because they require rewrites, not surface edits.

Can agents evaluate AI outputs more reliably than language models?

Eight-module agentic evaluation achieved 0.27% judge shift versus 31% for LLM-as-a-Judge on complex tasks. However, the memory module cascaded errors, revealing that agentic systems need error isolation mechanisms to maintain gains.

Show all 6 sources
Does user control over AI text shape feelings of ownership?

Study 1 found that greater user control over generated text raised sense of ownership, while personalizing the AI model had no impact on the AI Ghostwriter Effect.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.