Penalizing Transparency? How AI Disclosure and Author Demographics Shape Human and AI Judgments About Writing

Paper · arXiv 2507.01418 · Published July 2, 2025
Expertise in the Age of AI Content

As AI integrates in various types of human writing, calls for transparency around AI assistance are growing. However, if transparency operates on uneven ground and certain identity groups bear a heavier cost for being honest, then the burden of openness becomes asymmetrical. This study investigates how AI disclosure statement affects perceptions of writing quality, and whether these effects vary by the author’s race and gender. Through a large-scale controlled experiment, both human raters (n = 1,970) and LLM raters (n = 2,520) evaluated a single human-written news article while disclosure statements and author demographics were systematically varied.

This approach reflects how both human and algorithmic decisions now influence access to opportunities (e.g., hiring, promotion) and social recognition (e.g., content recommendation algorithms). We find that both human and LLM raters consistently penalize disclosed AI use. However, only LLM raters exhibit demographic interaction effects: they favor articles attributed to women or Black authors when no disclosure is present. But these advantages disappear when AI assistance is revealed. These findings illuminate the complex relationships between AI disclosure and author identity, highlighting disparities between machine and human evaluation patterns.

Introduction. AI systems are widely integrated into writing workflow as information sources [5, 21], readily available proofreaders [17], or “thought partners” [3, 7, 13, 22]. This transformation shifts our expectations surrounding authorship, originality, and intellectual labor [5, 16].

Readers, editors, and reviewers increasingly call for clear disclosure of AI involvement [4, 6]: Did the machine write this? How much of it? Can I still trust it? Beneath these questions lies a desire to calibrate judgments to discern the boundary between human insight and synthetic fluency.

Transparency, while noble in theory, is complicated in practice. Writers may hesitate to disclose their use of AI tools because they fear how their work will be perceived. Indeed, Baek et al. [2] found that labeling identical content as AI-assisted can reduce perceptions of credibility, creativity, and shareability. We wondered: Does this “AI discount” fall equally on all authors? Audiences do not assess writing in a vacuum; perceptions are shaped by the perceived identity of the author. For instance, Klaas & Boukes [14] found that male journalists were considered more credible than their female counterparts, particularly when reporting on technology.

Similarly, Eom et al.[8] showed that Black male scientists were perceived as warmer, while White female scientists were rated as more competent. If reader perceptions about AI disclosure are filtered through gendered, racialized, or other socio-epistemic cues, then the burden of openness becomes asymmetrical. Researchers conceptualize this negative societal judgments associated with someone’s perceived or actual use of AI systems as “perceptual harms.” [11] With these harms, the groups historically under-resourced and under-represented face the greatest risk of stigmatization for using AI.

Moreover, humans are no longer the sole evaluators of writing. Institutions are turning to AI to process large volumes of writing quickly, reduce costs, and sometimes, pursue the appearance of neutral judgment. In Texas, public school essays on standardized tests are now graded by AI systems trained to mirror human scorers, sparking concern among teachers and parents about fairness and accountability [12]. An estimated 99% of Fortune 500 companies use some form of automated screening in hiring [9]. AI-based performance review systems have been promoted to evaluate employee data in real-time and offer feedback based on actual metrics rather than human bias or memory [20]. In the publishing world, traditional publishers have started paying attention to metrics and even hiring data scientists to spot underrated gems that human editors might overlook [1]. Inkitt, a self-publishing platform, uses machine learning to evaluate user-submitted novels and identify breakout titles, guiding editorial investment and promotion [19].

These systems have become arbiters in high-stakes decisions, shaping whose writing is seen, trusted, or excluded.

This motivates us to investigate a question: When a writer discloses AI use, do readers judge the writing differently based

Method. 2 Human Participants Penalized AI-disclosed Articles across All Demographic Groups We conducted a pre-registered, large-scale survey with 1,970 participants.1 The study employed a 2×3×3 between-subjects factorial design with three independent variables: (1) AI disclosure (presence or absence of a disclosure statement), (2) author race (Asian, Black, White), and (3) author gender (man, woman, non-binary). Participants were randomly assigned to one of eighteen experimental conditions and evaluated an identical news article, with only the author biography and disclosure language varying across conditions.

The two types of disclosure statements used in the study are shown in Table 1. Additional details on the rationale behind this research design are provided in Appendix B.

Table 1. Disclosure statements shown to participants.

Condition Disclosure Statement Control Statistical information updated as of Oct. 11, 2024. AI This article was created with assistance from Artificial Intelligence (AI) tools. Statistical information updated as of Oct. 11, 2024.

The study has four phases. In Phase 1, participants were exposed to identical, human-authored news article across all conditions, with variations in the author biography, photo, and disclosure statement (See Appendix A). The control condition includes a basic statistical update disclosure, while the treatment condition explicitly states the use of AI tools in content creation. During Phase 2, participants engaged in a deception task where they identify evidence and recommendations from each paragraph. Phase 3 consisted of dependent measures where participants rate four key aspects using 7-point Likert scales (1 = Strongly Disagree, 7 = Strongly Agree): information trustworthiness, comprehensiveness, writing quality, and likelihood of sharing. In Phase 4, participants filled an exit questionnaire capturing recall of the disclosure statement and author demographics, personal and professional AI usage patterns.

To examine whether author identity moderated the AI disclosure effect, we fit a series of linear models with interaction terms. In

Discussion. 4 Disclosure, Author’s Identity, and the Fragility of LLM Alignment Our two-pronged investigation shows that humans and LLM raters alike apply a modest but consistent penalty to articles when AI assistance is disclosed. This finding shows that AI disclosure carries epistemic stigmatization, although the magnitude of “AI disclosure discount” is small—less than 0.15 on a 7-point scale across all experiments, indicating a perceptible yet not overwhelming penalty. This finding aligns with prior work showing that AI disclosure has a statistically significant but relatively modest impact on perception [18].

However, a key divergence emerged: interaction effects between author identity and disclosure status were present only in LLM evaluations. Among human raters, disclosure penalties were relatively uniform across authors of different races and genders. In contrast, both LLMs exhibited identity-disclosure interaction effects. GPT-4o-mini showed a pronounced preference for Black authors in the control condition, which diminished when AI involvement was disclosed. Qwen2.5-7B-Instruct similarly favored woman authors in the absence of disclosure, a gender bias that disappeared when AI assistance was acknowledged.

Notably, both favored groups—Black authors and women—are historically marginalized. These preferences may reflect an alignmentdriven over-correction, in which models trained with human feedback disproportionately reward underrepresented identities [15].

However, when AI disclosure is present, these fairness-oriented preferences vanish. We term this dynamic a form of “vanishing alignment,” where the social and ethical calibration of model behavior becomes fragile under changing contextual cues. A similar phenomenon has been observed by Hofmann et al. [10], who found that LLMs trained with human preference alignment projected overtly positive stereotypes toward African American English speakers while maintaining covertly negative biases that surfaced in indirect prompts. Our findings suggest that similarly, fairness-aligned behaviors in LLMs may be conditional and unstable.

Several factors may help explain why these effects appear in LLMs but not in human ratings. While human raters bring diverse lived experiences and norms to the task, LLM raters might produce more stable, patterned outputs across multiple runs, making their biases more legible. Moreover, the genre of the writing (news articles) may also play a role [23]. AI disclosure may trigger genre expectations around objectivity, amplifying penalties for AI involvement and reducing the salience of the author’s race or gender.

Conclusion. Our research offers groundwork for understanding how AI disclosure shapes perceptions of writing and author’s identity. Future work should explore the epistemic consequences of AI disclosure across genres, audiences, and forms of AI involvement and critically, attend to how perceptual harms may accumulate or compound in various evaluative settings.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do AI hiring systems affect authenticity, fairness, and candidate preferences? Why do confident AI outputs mislead human trust calibration? Can base models hide emergent misalignment through alignment training? How does AI-generated content create social proof without authentic interaction? Does disclosing AI authorship change how audiences evaluate the writing? How can we detect and account for LLM involvement in academic writing? Do language models reason through disagreement or only accommodate it? How do writers navigate authorship and delegation with AI? How reliably can humans and AI detectors identify machine-generated text? Can persona profiles improve LLM prediction accuracy and consistency?