A light-touch AI literacy intervention helps protect against AI political persuasion

Paper · arXiv 2609.16432 · Published September 14, 2026
Argumentation and Persuasion

Conversations with large language models (LLMs) can substantially shift beliefs and attitudes, raising concerns about manipulation using AI persuasion. Here we test whether a light-touch AI literacy intervention – a brief warning that LLMs can be prompted to persuade and may present information selectively – helps protect users. Across two experiments (total N = 3,208 Americans) in which participants conversed with an LLM instructed to shift their views about different political topics, the presence of a warning reduced belief change by roughly one-half (-48.1%, 95% CI [-59.5%, - 36.8%]) relative to the control. Importantly, the warning did not significantly reduce trust in generative AI more broadly. Light-touch literacy interventions can help protect users against AI political persuasion.

Introduction. There is increasing evidence that conversations with AI chatbots can be highly persuasive. While these conversations can be used in beneficial ways, such as debunking conspiracy theories (Costello et al., 2024) or reducing science skepticism (Hornsey et al., 2026), there is substantial concern about AI dialogues being used to persuade in a harmful manner, such as persuading voters about political issues (Hackenburg et al., 2025; Salvi et al., 2025; Argyle et al., 2025) and candidates (Lin et al., 2025; Potter et al., 2024), and eroding democratic norms (Schroeder et al., 2025). Despite these concerns, little work has developed or tested solutions to protect users from influence by conversational AI. Here we ask whether a minimal AI literacy intervention can reduce susceptibility. In two studies, we test the effect of informing participants about the potential for large language models (LLMs) to be prompted to persuade, and thus to provide biased or selective information (see Fig. 1A).

Discussion / Conclusion. Given the evidence that AI chatbots can persuade across a wide range of issues, there are widespread calls to find ways to limit these persuasive effects. Here, we present evidence that a light-touch AI literacy intervention – simply informing people that AI models may have motives to persuade or manipulate (even without providing specific information about the model’s intent) – can reduce persuasion in political settings by approximately one-half. Importantly, this intervention did not have a significant effect on trust in generative AI more broadly. This suggests that the warning is specifically conferring protection against political persuasion. While the intervention does not entirely eliminate AI’s persuasive effects, it is a proof of concept that literacy treatments can have a meaningful impact. Future work should establish how to most effectively deliver such information. It is also important to further investigate ways to reduce the influence of manipulative AI while preserving the benefit of accurate AI (making people more discerning rather than more generally skeptical; Guay et al., 2023). The lack of effect on overall trust in generative AI indicates that our treatment was targeted at least to some extent; future work should explore effects on prosocial persuasion.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do false presuppositions and sycophancy drive persistent false beliefs in models? What factors drive AI persuasiveness and how can it be mitigated? Do language models reason like humans or mimic surface patterns? Do writers recognize when AI writing assistance alters their expressed stance? Why do people disclose to AI systems despite their artificial nature? How can AI chatbots provide therapeutic benefit without causing harm? How do neighboring agents influence whether others cooperate or collude? How well do AI systems understand human social norms? Should AI communication design follow human conversation norms or develop distinct machine-specific principles? Can multi-agent systems avoid converging on false agreement without deliberation? Does transformer attention architecture inherently drive sycophancy?