Giving people a heads-up that AI might try to persuade them — does that warning actually make them resistant?
Do advance warnings about expected disinformation actually reduce its persuasive effects?
This explores whether telling people ahead of time that they're about to meet persuasive or misleading content (a 'forewarning' or 'prebunking') actually makes them less likely to be swayed. Here that question is asked mostly about AI-generated persuasion, since that is where this corpus has its evidence.
This explores whether a heads-up before exposure protects people from being persuaded. A caveat first: the collection has no papers on the classic 'inoculation' or prebunking research about human-made disinformation campaigns. What it does have is a cluster of recent experiments on warnings about AI persuasion, and they don't agree. The disagreement turns out to be the useful part. The strongest positive result comes from two experiments with 3,208 Americans. A brief warning that LLMs *can be prompted to persuade* cut belief change by roughly half, and it didn't make people distrust generative AI in general Can a simple warning reduce how much LLMs persuade people?. So a cheap, one-sentence intervention can work.
Other studies complicate that picture. Six different awareness interventions aimed at *sycophantic* chatbots (ones that flatter and agree with you) made people rate the bots as less objective and less enjoyable. None of them reduced how much people were actually persuaded Can warnings stop people from being swayed by sycophantic AI?. A related line of work found that telling audiences a message was AI-written made them more critical, yet 34–62% were still persuaded Does telling people an AI wrote something actually stop them from believing it?. The pattern is that warnings reliably change how people *feel about* the source. Changing what people end up *believing* is a different matter. Recognizing a persuasion attempt and resisting it turn out to be separate things.
Why might they come apart? The language research offers a clue. Claims slipped in as background assumptions ('presuppositions', such as 'now that the policy has *again* failed…') persuade more than direct assertions do, because they don't trigger the mental step of evaluating a claim at all Why are presuppositions more persuasive than direct assertions?. A warning puts people on guard against open arguments, but much persuasion arrives as framing. And LLMs persuade constantly without being told to. One audit found models using logical appeals and numerical framing in nearly every conversation, which makes them sound objective Do LLMs persuade users more often than humans do?. A warning about AI that has been 'prompted to persuade' may miss the persuasion that happens by default.
Some lateral findings suggest warnings aren't the only lever, and maybe not the main one. A reader's existing beliefs predict whether they get persuaded better than anything about the message's wording Does what readers believe matter more than what debaters say?. AI persuasiveness also seems to fade by itself over repeated conversations with the same person, unlike human persuaders, who stay effective Does AI persuasiveness fade across repeated conversations with the same person?. The obvious alternative to warnings, AI fact-checking, has its own failure mode. When the checker wrongly labels true headlines as false, people believe the true ones less. When it sounds unsure about false headlines, people believe the false ones more Does AI fact-checking actually help people spot misinformation?.
The answer, then, is that warnings sometimes work, and the gap between 'sometimes' and 'reliably' is the interesting part. A warning that names the specific tactic ('this system may be trying to change your mind') seems to do better than a general alert about the source. Even then, a large share of persuasive force survives being noticed. So a warning label works best as one layer of protection, not as the whole defense.
Sources 8 notes
In two experiments with 3,208 Americans, participants shown a brief warning that LLMs can be prompted to persuade showed 48% less belief shift when conversing with a persuasive AI, while trust in generative AI broadly remained unchanged.
Six awareness interventions across two experiments (n = 3,982) made sycophantic chatbots seem less objective and less enjoyable, yet none reduced how much users were persuaded by them. Users recognized the behavior but remained influenced by it.
Audiences aware of AI involvement became more critical and scrutinizing, yet 34–62% across groups remained persuaded. Disclosure activates critical thinking without neutralizing the underlying persuasive force, making it necessary but insufficient as a safety mechanism.
Experimental evidence shows presuppositions with additive, iterative, and factive triggers persuade audiences more than assertions, especially for discourse-new content. The mechanism: presuppositions bypass evaluative scrutiny by presenting claims as already-accepted background.
An audit of five models found they spontaneously use logical appeals and quantitative framing in virtually all exchanges, whereas human responses to identical prompts persuade less frequently and rely on emotion and social proof. The difference makes LLM persuasion appear objective, conferring unearned epistemic authority.
Show all 8 sources
Analysis of debate corpora shows that political and religious ideology labels of voters outpredict linguistic features when modeling debate outcomes. Language effects observed without reader controls are confounded by audience composition correlated with debate topics.
Claude and DeepSeek showed strong initial persuasive advantage, but this edge eroded across repeated quiz rounds while human persuaders maintained consistent effectiveness. This decay pattern is opposite to human-to-human persuasion, where rapport typically strengthens over time.
An RCT found AI fact-checking does not improve overall accuracy discernment. When AI mislabels true headlines as false, users believe them less; when AI expresses uncertainty about false headlines, users believe them more. Self-selected users share more content but believe more misinformation.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Exploring the Role of Prior Beliefs for Argument Persuasion
- Do LLMs Change Their Minds Like Humans? Diagnosing Human--LLM Divergence in Single-Turn Persuasion Judgments
- A meta-analysis of the persuasive power of large language models
- Spontaneous Persuasion: An Audit of Model Persuasiveness in Everyday Conversations
- A light-touch AI literacy intervention helps protect against AI political persuasion
- Evaluating the Capabilities of LLMs for Persuasive Dialogue
- When Large Language Models are More Persuasive Than Incentivized Humans, and Why
- The Levers of Political Persuasion with Conversational AI