SYNTHESIS NOTE
Topics›Human Centered Design›this note

Can we distinguish helpful explanations from manipulative ones?

Rhetorical strategies used to justify appropriate AI adoption rely on the same persuasion mechanisms as dark patterns. Without observable intent, explanation and manipulation look identical—raising urgent questions about how to audit XAI systems responsibly.

Synthesis note · 2026-05-02 · sourced from Human Centered Design
How do people decide what to share with AI systems?

The Rhetorical XAI paper acknowledges the structural tension at the heart of its own framework. Citing Gray et al. on dark patterns and Chromik et al.'s extension of dark patterns to XAI, it notes that the same rhetorical machinery used to communicate why AI merits appropriate use can be deliberately deployed to exploit cognitive and emotional vulnerability and steer users toward unintended decisions. There is no clean separation between rhetorical XAI for appropriate adoption and rhetorical XAI for coercion. Logos, ethos, and pathos are channels, not intentions; the same persuasive load can recruit cooperation or extract compliance, and the artifact-level signature is identical.

This is not a marginal concern, it is a structural one. If explanation effectiveness depends on rhetorical work, and rhetorical work is the same set of mechanisms used in dark patterns, then the audit problem becomes severe: the explanation that responsibly justifies adoption looks, from the outside, like the explanation that manipulates. Effectiveness metrics that reward "users acted on the explanation" cannot distinguish appropriate adoption from successful coercion. The distinction lives in the designer's intent and the user's actual interest, neither of which is recoverable from the artifact in isolation.

This is a related-risk pair to Does polished AI output trick audiences into trusting it? — both insights describe how persuasive surface form does work that should be done at a different layer (deliberation, expert judgment) without that layer being visible. It also connects to Do people prefer AI moral reasoning when they don't know the source?: when AI authorship is hidden, persuasion lands; when revealed, it is rejected. Disclosure interacts with rhetorical effectiveness in a way that any responsible XAI deployment has to specify. Hidden rhetorical work is dark by default, even when intentions are clean.

For the False Punditry / Knowledge Custodian writing thread, this is the structural form of the concern. The same explanation that helps a user calibrate trust can be tuned, with no change in form, to over-extract trust. Calling rhetorical XAI "explanation" is itself a rhetorical choice that obscures this — and the field has not yet developed evaluation criteria that hold across the appropriate-adoption / coercion gap.

Inquiring lines that read this note 44

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Does augmenting symbolic reasoning improve LLM logical reasoning ability? What determines AI's persuasive power and how can it be detected or mitigated? How can humans maintain effective oversight as AI systems scale? How do users confuse explanation quality with actual system accuracy? Does disclosing AI authorship change how audiences evaluate the writing? Can AI systems participate in genuine communication or only simulate it? How should AI agents balance proactive engagement with conversational respect? Why do models reveal hidden associations despite concealment attempts? What gaps exist between benchmark performance and real deployment outcomes? How do interpretive frames override surface features in text comprehension? Can humans reliably detect and resist AI-generated misinformation? Why do language models struggle to implement user intent accurately from prompts? How can emotionally responsive AI maintain reliability and healthy boundaries? Why does polished AI output gain credibility despite fundamental verifiability problems? How do educators verify student capability when AI can produce indistinguishable work? Does chain-of-thought reasoning reveal how models actually think or merely imitate reasoning? Why do confident AI outputs mislead human trust calibration? Can monitoring reasoning traces and behavior detect hidden agent deception? Can AI systems evade safety evaluations through reasoning manipulation? How do AI systems determine and balance multiple competing objectives? Are AI-generated articles systematically disadvantaged in search ranking and user engagement?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 127 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

rhetorical strategies shade into dark patterns — the same persuasion mechanisms that justify appropriate adoption can manipulate cognitive and emotional vulnerability