Do LLM raters show hidden demographic preferences that disclosure erases?
Explores whether language models systematically favor certain demographic groups when their AI involvement is not disclosed, and whether that preference disappears under transparency. This matters because it reveals potential fragility in AI alignment training.
Only the LLM raters in this study show demographic interaction effects. Human disclosure penalties were "relatively uniform across authors of different races and genders." The two LLMs were not uniform. GPT-4o-mini "showed a pronounced preference for Black authors in the control condition, which diminished when AI involvement was disclosed." Qwen2.5-7B-Instruct "similarly favored woman authors in the absence of disclosure, a gender bias that disappeared when AI assistance was acknowledged." The authors note that both favored groups "are historically marginalized." They name the pattern "vanishing alignment," described as a case "where the social and ethical calibration of model behavior becomes fragile under changing contextual cues."
The authors' explanation is hedged. The preferences "may reflect an alignment-driven over-correction, in which models trained with human feedback disproportionately reward underrepresented identities." They tie the pattern to Hofmann et al., who found that preference-aligned LLMs projected "overtly positive stereotypes" toward African American English speakers "while maintaining covertly negative biases that surfaced in indirect prompts." For why the effect appears in LLMs and not in humans, they offer two conjectures. LLM raters "might produce more stable, patterned outputs across multiple runs, making their biases more legible." And news-genre expectations of objectivity may amplify penalties for AI involvement, "reducing the salience of the author's race or gender." Neither conjecture is tested in the excerpt.
Set against the nearest notes, this note moves the question from the text to the judge. Does AI writing assistance change how readers perceive the writer? reports that AI assistance distorts demographic markers in what readers see. Here the concern is that the scoring models' own demographic preferences depend on a contextual cue. The cue is the same kind of lever as in Does telling people an AI wrote something actually stop them from believing it?, which found disclosure raises scrutiny without collapsing persuasion. In this study the same cue switches off a demographic preference in two LLMs. The human sample does not share that pattern, since its disclosure penalty is uniform across groups.
The excerpt does not establish the size of the demographic effects. It says the authors fit "linear models with interaction terms" but reports no estimates, test statistics, or number of LLM ratings per condition. The abstract refers to "LLM raters" generally, while the discussion names only two models, so the excerpt does not show that they are the full set tested. The over-correction account is a possibility the authors raise, not a finding. The implication is conditional. If the pattern holds beyond these two models, an LLM used to score writing would need audits that vary contextual cues such as disclosure, not only the author's identity. On this excerpt alone, the claim is that two models showed the pattern under one news article and one disclosure wording.
Inquiring lines that read this note 25
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do AI hiring systems affect authenticity, fairness, and candidate preferences?- Does recruiter use of generative AI change how they evaluate AI skills in candidates?
- Do recruiters understand what their hiring algorithms actually prioritize?
- Do candidates prefer being screened by AI or by humans?
- Can employers tell when applicants use generative AI tools?
- Can existing fairness audits detect LLM self-preference in hiring systems?
- Would human recruiters supervised by AI show similar self-preference patterns?
- Can simple interventions like system prompting reduce LLM self-preference in hiring?
- Does user preference for AI suggestions encode cultural reliance gaps?
- Do younger voters and voters of color trust AI differently?
- Does the disclosure penalty vary based on article genre or topic?
- How do cultural backgrounds shape reactions to disclosed AI authorship?
- Why does the disclosure penalty still hold even when readers have high AI literacy?
- Does directly adopted AI text require disclosure even if methodology is unchanged?
- Why does disclosure of LLM authorship change reader trust and preference?
- Does LLM use reduce writing costs differently across linguistic backgrounds?
- How does stylistic matching contribute to LLM self-preference in evaluations?
- Can reviewer-author matching by LLM use amplify biases in acceptance decisions?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does telling people an AI wrote something actually stop them from believing it?
When audiences learn that AI created content, do they become skeptical enough to resist its persuasive pull? This explores whether disclosure works as a genuine defense against AI-driven persuasion or merely shifts how people process it.
the same disclosure lever: there it raises scrutiny of persuasion, here it removes a demographic preference in LLM raters.
-
Does AI writing assistance change how readers perceive the writer?
Explores whether AI-assisted writing systematically alters reader impressions of the writer's political views, competence, emotion, and demographic identity. Understanding this matters because perception shapes trust and influence in public discourse.
that note measures demographic distortion in the text; this one finds the judging models' demographic preferences depend on a cue.
-
Does disclosing AI assistance make readers trust articles less?
When articles carry a label saying they used AI tools, do human and AI raters downgrade their quality assessments? This matters because writers worry disclosure could harm how their work is received.
sibling note: the shared penalty against which the LLM-only demographic pattern is set.
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Penalizing Transparency? How AI Disclosure and Author Demographics Shape Human and AI Judgments About Writing
- AI Self-preferencing in Algorithmic Hiring: Empirical Evidence and Insights
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
- Unintended Impacts of LLM Alignment on Global Representation
- LLM Evaluators Recognize and Favor Their Own Generations
- AI Suggestions Homogenize Writing Toward Western Styles and Diminish Cultural Nuances
- Scientific production in the era of Large Language Models
Original note title
LLM raters favor Black or women authors when AI use goes undisclosed, and the preference vanishes once it is revealed