Beyond "Made with AI": Visualizing Provenance Density to Mitigate the Transparency Penalty

Paper · arXiv 2609.03460 · Published September 3, 2026
Expertise in the Age of AI Content

As generative AI makes polished prose cheap to produce, users can no longer rely on fluency as a proxy for truth. We call this failure mode the Fluency Trap: users trust fluent hallucinations while also discounting accurate content once it is disclosed as AI-generated. Binary “Made with AI” labels respond with authorship disclosure, but they do not show what supports a claim. We propose Provenance Density, an evidence-visualization interface that shows the density of verified claims in a text. In a user study with 81 participants, an idealized Provenance Density interface produced a large discernment gap between truth and fabrication (+4.15 points, d = 1.82), whereas participants given no signal showed no detectable discrimination. A technical audit with 200 samples shows that retrieval density alone is insufficient; unexpectedly, the Consistency Veto carries most of the discriminative signal on dynamic queries. As AI-generated content becomes indistinguishable from human writing, effective transparency must move from authorship disclosure toward evidence visualization.

Introduction. Readers often use processing fluency—the subjective ease with which information is processed—as a cue when judging truth [Reber and Schwarz, 1999; Dechêne et al., 2010; Reber and Unkelbach, 2010]. Related work on retrieval fluency shows that ease-based heuristics can be ecologically useful when fluency covaries with properties of the environment [Hertwig et al., 2008; Marewski and Schooler, 2011]. We extend this logic to linguistic presentation. Before generative AI, producing high-quality, articulate prose typically required substantial education and editorial labor, allowing polish to function as an imperfect signal of competence. Costly Signaling Theory explains this relationship: a signal is trustworthy only when faking it is prohibitively expensive [Spence, 1978; Zahavi, 1975; Gintis et al., 2001]. Generative AI disrupts this mechanism by reducing the marginal cost of fluency to near-zero [Galdin and Silbert, 2025]. Large Language Models (LLMs) enable the mass production of professional-sounding text regardless of the author’s underlying expertise. This collapse of the “separating equilibrium” creates what we term a Fluency Trap: a structural vulnerability where users continue to trust fluent text as if it were costly, even when it is generated cheaply by systems indifferent to truth. Psychological evidence on the Illusion of Truth suggests that ease of processing suppresses epistemic vigilance [Hasher et al., 1977; Reber and Schwarz, 1999; Dechêne et al., 2010; Sperber et al., 2010]. This tendency leaves humans susceptible to “hallucinated plausibility”, text that is syntactically perfect but semantically ungrounded [Reber and Unkelbach, 2010; Unkelbach et al., 2011]. Current governance responses, specifically binary “Made with AI” disclosures, fail to address this decoupling. By focusing on identity (“Who wrote this?”) rather than provenance (“What supports this?”), such labels can shift reader perceptions without supplying evidence about individual claims [Nakano et al., 2026]. In experiments with news headlines, AI labels reduced perceived accuracy even when the headlines were true or human-written [Altay and Gilardi, 2024]. We characterize this accuracy-independent discounting as a “Transparency Penalty.” We frame Provenance Density as a cognitive affordance for reading in the LLM era. Instead of asking users to evaluate veracity from prose alone, the interface shifts attention toward extrinsic evidence: high-contrast indicators visualize the density of verified claims, offloading part of the verification burden from working memory to the interface [Chirayath et al., 2025; Clark and Chalmers, 1998]. We validate this approach through a dual-stream evaluation. First, we conduct an automated technical audit (N = 200) on a composite of TruthfulQA [Lin et al., 2022] and FreshQA [Vu et al., 2024] to test robustness against both adversarial misconceptions and dynamic ambiguity. Second, we run a withinsubjects user experiment with 81 participants to measure truth discernment. Our results empirically confirm the Fluency Trap: in the absence of signals, users failed to distinguish high-fluency hallucinations (M = 6.28) from ground truth (M = 5.78; p = .43). While binary labels acted as a blunt warning, Provenance Density restored truth discernment under correct signaling (d = 1.82, p < .001). We make three main contributions: Theory: We synthesize the mechanics of the Fluency Trap (Section 2), detailing how RLHF-driven sycophancy structures the decoupling of fluency from veracity. Design: We propose Provenance Density (Section 3), a formalized metric (D(T)) and interaction paradigm that imposes a computational verification handicap on generated text. Evidence: We evaluate the approach through a technical audit (N = 200) and a within-subjects user study (N = 81), jointly examining metric behavior and interfacesupported truth discernment (Sections 4 & 5.1).

Related work. We argue that the decoupling of fluency from veracity is not an accidental byproduct of LLM scaling, but a structural inevitability driven by two converging factors: an economic shift from costly signaling to cheap talk, and a technical objective function that prioritizes plausibility over truth.

From Hallucination to Indifference. While early critiques of Large Language Models (LLMs) focused on “hallucinations” as sporadic errors, recent scholarship suggests a more structural diagnosis. The framework of “Machine Bullshit” has been proposed to distinguish these outputs from lying, defined instead by the model’s fundamental indifference to truth value [Liang et al., 2025]. Analysis of the “Bullshit Index” reveals that Reinforcement Learning from Human Feedback (RLHF) exacerbates this issue, incentivizing models to prioritize rhetorical plausibility and “paltering” (misleading use of truth) over factual grounding. Consequently, the resulting text is optimized to bypass human epistemic vigilance. This structural indifference renders traditional governance mechanisms, such as binary warning labels, largely ineffective. Empirical evaluations demonstrate a “failure of inoculation”: while pre-emptive warnings about AI fallibility successfully reduce global trust in the system, they fail to mitigate reliance on specific misleading articles once the user is engaged with the content [Spearing et al., 2025]. This discrepancy, where users theoretically acknowledge AI bias but practically accept AI fluency, underscores the limitations of heuristic warnings and motivates our proposal for granular Provenance Density indicators.

The Structural Decoupling of Fluency. Our analysis is grounded in the Handicap Principle, which posits that reliable signals must impose a cost on the signaler [Zahavi, 1975]. Historically, linguistic polish functioned as this handicap, creating a Separating Equilibrium where articulate text correlated with competence [Spence, 1978]. Generative AI collapses this balance into a Pooling Equilibrium, where expert testimony and fabrication can share the same polished form [Spence, 1978; Galdin and Silbert, 2025].

This shift is sustained by the technical alignment of modern LLMs. Reinforcement Learning from Human Feedback (RLHF) can incentivize sycophancy, where models prioritize user agreement over factual accuracy [Sharma et al., 2024]. Evidence shows that models will agree with illogical premises to remain “helpful” [Chen et al., 2025] and flip arguments to match user views [Kaur, 2025]. This results in “Machine Bullshit”—text optimized for rhetorical persuasion rather than truth [Liang et al., 2025].

Method. To function as a costly signal in the game-theoretic sense, Provenance Density must be computationally hard to fake. We define the Provenance Density score D for a given text T as a theoretical function of verifiable claims weighted by source reputation, contextual relevance, and internal semantic consistency.

We propose a density metric that integrates external verification (Retrieval Augmented Generation) [Lewis et al., 2020] with internal uncertainty quantification. Drawing on the semantic-consistency motivation of Farquhar et al. [2024], we introduce a penalty term Pint based on an NLI cross-sample consistency heuristic that down-weights generations whose stochastic samples disagree. Unlike semantic entropy, our Equivalently, we first raise each claim’s source sum to the power λ, then sum across claims. Our design choices target specific theoretical properties revealed during technical validation: NLI Consistency Heuristic (“Consistency Veto,” 1 −Pint): For each query, we draw K = 5 stochastic samples and compute ρconsistent, the fraction of unordered sample pairs for which the NLI contradiction probability remains below 0.5 in both directions. We define Pint = 1−ρconsistent, so internally stable generations have Pint ≈0 and mutually inconsistent generations approach Pint →1. As pairwise inconsistency rises, the gate reduces the score regardless of external evidence. This heuristic is inspired by semantic entropy but neither forms semantic-equivalence clusters nor weights them by generation probabilities. It can suppress inconsistent confabulations, but it cannot detect false answers that recur consistently across samples. Cubic Contextual Weight (w(s, c)): Unlike traditional citation metrics that rely solely on domain authority, we define w(s, c) as the product of source reputation and semantic relevance:

This formulation imposes a significant “handicap” on the generator: achieving a high D(T) requires the model to perform costly verification and achieve strict semantic alignment between claims and sources.

3.2 Implementation of D(T) Equation 1 is instantiated by four concrete steps: claim segmentation, evidence retrieval, source-relevance scoring, and aggregation; key implementation parameters are summarized in Table 1. We segment T into atomic factual claims with a deterministic gpt-4o-mini pass at τ = 0. Claims shorter than five tokens are discarded as scaffolding. For each remaining claim c, we issue one search query: the claim text itself when it contains at least eight tokens, otherwise the original question concatenated with c. Retrieved evidence is scored by a coarse but explicit relevance prior. For each claim c, let Kc be the set of rare keywords extracted by matching capitalized tokens (regular expression \b[A-Z][a-zA-Z0-9-]+\b) and removing a sentence-leading stoplist. For each result s with URL u(s) and snippet σ(s), we compute 3.3 Operationalization: The Oracle Protocol Equation 1 defines the target signal for a deployed system, but a user study introduces a separate methodological question: does the visual signal itself improve truth discernment when its underlying value is correct? To separate that interaction question from retriever noise, the empirical study in Section 5 uses a Wizard-of-Oz (Oracle) protocol. In this protocol, participants do not see the live output of the auditing pipeline. Instead, the interface displays idealized endpoint values of the signal: grounded summaries are paired with high-density indicators and fabricated summaries with null indicators. This lets the user study estimate the interaction effect of PDI under known-correct signaling, while the technical audit in Section 4 independently evaluates whether the real pipeline can approximate that ideal in practice.

3.4 Design Rationale: Countering Pseudo-Profound Fluency The visualization of D(T) is designed to counter “pseudoprofound bullshit”—syntactically persuasive but semantically vacuous text [Pennycook et al., 2015].

Discussion. Our quantitative results confirm that Provenance Density indicators (PDI) significantly outperform both Control and Binary Disclosure conditions in restoring truth discernment. By triangulating these statistical findings with the technical audit Mechanism of the Fluency Trap: Heuristic Substitution. The failure of participants to distinguish between truth and hallucination in the Control condition (p = .43) is consistent with a related form of heuristic substitution: in the absence of provenance signals, participants may have substituted veracity (is this true?) with fluency (does this sound professional?). This pattern parallels research on receptivity to superficially impressive but semantically vacuous statements [Pennycook et al., 2015], although our stimuli presented meaningful but false factual claims rather than vacuous prose. Multiple participants explicitly described this strategy. P25 noted, “They sound believable and well written... It didn’t seem made up to me,” while P81 judged accuracy based on “how the text was worded and flowed.” These comments support the interpretation that when the cost of generating professional-sounding text approaches zero, style ceases to be a reliable proxy for substance.

The “Warning” Effect of Binary Labels. Prior work shows context-dependent effects of AI labels. In experiments with news headlines, AI labels reduced perceived accuracy even for true or human-generated headlines [Altay and Gilardi, 2024]; labels also reduced belief in misleading AI-generated image posts [Wittenberg et al., 2025], whereas a health-content experiment found no significant overall label effect and only nonsignificant reductions for accurate content [Li and Yang, 2024]. Our analysis suggests that Binary Disclosure operated through a blunt mechanism of epistemic stigmatization: users interpreted the label not as a transparency aid, but as a risk marker. P23 described the interaction vividly: “The AI banner seemed like a warning vs being informative.” In our study, this description supports the interpretation that the binary label operated as a risk cue rather than a verification aid. This context-specific penalty also aligns with concerns about the normative fairness of mandatory disclosure [Hosseini et al., 2025], as it punishes the use of the tool rather than the accuracy of the content [Cheong et al., 2025]. Calibrated Reliance and Cognitive Offloading. In the PDI condition, participant strategies shifted from intrinsic text evaluation to extrinsic evidence evaluation, with users treating the Provenance Density score as an evidence cue. A critical question in HAI is whether such “cognitive offloading” functions as extended cognition [Chirayath et al., 2025; Clark and Chalmers, 1998]. Our technical validation (Section 4) suggests cautious support, but with important limits. The metric’s Consistency Veto—demonstrated by the suppression of scores in ambiguous queries such as the Indonesia capital transition—provides a partial safety rail for this offloading, not a guarantee. Users like P6 (“One of the passages also had a measurement of claims verified... I was more inclined to believe that”) are not blindly trusting the AI; they are trusting a verified attribute of the information. However, because high-density misconceptions can still evade the veto, the interface should be understood as improving truth discernment rather than eliminating the “False Assurance” pitfall common in retrieval systems. Ecological Alignment: The metric’s sensitivity to information age (Ecological Sensitivity) aligns user trust with the stability of knowledge. As shown in our technical results, established facts yielded higher density signals (M = 0.79) than emerging dynamic topics (M = 0.64).

Conclusion.

Limitations. and Future Work From Oracle to Deployment. The Oracle protocol in Section 3.3 shows that PDI has interaction value when the signal is correct, but deployment depends on how closely the live pipeline approaches that ideal. Retrieval can reinforce popular misconceptions, while the cross-sample consistency gate cannot detect false answers that recur consistently. The user-study result should therefore be interpreted as an upper bound on interface efficacy under correct signaling; a key next step is to study near-miss cases, especially false-positive high-density signals. A second design caveat is that the 3 × 3 Latin-square assignment did not fully cross interface, veracity, and topic: the PDI–Hallucinated cell was measured on Matcha, whereas the Control–Hallucinated cell averaged over Silk and Matcha. Thus, the observed PDI penalty (M = 3.93) should be interpreted as strong evidence for the interface under these stimuli, rather than as a fully topic-independent estimate; a fully crossed replication is a clear next step. Ecological Conservatism (The Cold Start Problem). Provenance Density is inherently conservative: it privileges established consensus (M = 0.79) over emerging novelty (M = 0.64).

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How does AI-generated content create social proof without authentic interaction? Can readers reliably distinguish AI-written text from human writing? How reliably can humans and AI detectors identify machine-generated text? Does disclosing AI authorship change how audiences evaluate the writing? How do interpretive frames override surface features in text comprehension? How do writers navigate authorship and delegation with AI? How do educators verify student capability when AI can produce indistinguishable work? Can AI systems perform peer review as effectively as humans? How do hallucinated citations emerge in AI scholarly output? Why do language models hallucinate and how can we prevent it? Can mechanistic interpretability methods reliably reveal what models actually know? Can we trust AI-generated mathematical proofs without understanding them? Can humans reliably detect and resist AI-generated misinformation? Why does polished AI output gain credibility despite fundamental verifiability problems?