Do LLMs Change Their Minds Like Humans? Diagnosing Human--LLM Divergence in Single-Turn Persuasion Judgments

Paper · arXiv 2608.29803 · Published August 30, 2026
Argumentation and Persuasion

Large language models (LLMs) are increasingly deployed as proxies for human participants in social simulations, yet whether they update their beliefs in response to persuasive arguments, as humans do, remains poorly understood. We conduct a systematic comparison using a naturally occurring online persuasion corpus in which original posters explicitly verify whether a reply changed their view. Our results show that LLMs achieve only slight agreement with humans (Cohen’s κ ranging from 0.079 to 0.178). Content-level analyses show that humans and LLMs agree on the strongest persuasion cues but diverge on finer ones: humans are more swayed by novel content and assertive language, whereas LLMs favor topical similarity and surface-level formatting. At the level of persuasion strategy, LLMs underweight emotional appeals and overweight credibility signals relative to humans, while the type of proposition under debate exerts no measurable effect on the degree of divergence. Furthermore, switching from first-person role-playing to third-person observation shifts all models toward greater resistance to persuasion, with the effect varying across persuasion strategies and textual features.

Introduction. Large language models (LLMs) are increasingly used for simulating human interactions in various contexts (Park et al., 2023; Argyle et al., 2023; Gao et al., 2024), including online discourse (Chuang et al., 2024), political elections (Zhang et al., 2024), and collective decision-making (Jarrett et al., 2025). A key cognitive process in these interactions is be- lief updating (Anderson, 1981; Hogarth and Einhorn, 1992), through which individuals selectively revise their prior beliefs after encountering new evidence or arguments. When exposed to the same arguments that humans find persuasive or unpersuasive, do LLMs revise or maintain their positions in a similar manner? If systematic divergence exists, applications that assume human-like reasoning risk producing distorted outcomes (Chen et al., 2026; Anthis et al., 2025). Therefore, to achieve simulations with fidelity, it is critical to understand whether such divergence exists, and if so, what modulates it.

Discussion / Conclusion. In this study, we systematically compare LLM belief update judgments against human-verified persuasion outcomes. All tested models achieve only slight agreement with human labels. The overall rate of divergence is similar across models, but its internal composition differs markedly. Diagnosing what drives this divergence, we find that it is insensitive to the type of claim under debate but systematically shaped by how arguments are constructed. LLMs and humans rely on qualitatively different cues when evaluating persuasive force, with LLMs favoring topical overlap and credibility signals while discounting novelty and emotional engagement. Introducing a third-person observer perspective shifts all models toward greater resistance to persuasion, but the effect is uneven across persuasion strategies and textual features. These patterns point to a structural mismatch between LLM and human belief updating. Such a mismatch persists even in the strongest model tested, and thus whether further scaling resolves it remains an empirical question for future analysis.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Can multi-agent systems avoid converging on false agreement without deliberation? What factors drive AI persuasiveness and how can it be mitigated? Do language models reason like humans or mimic surface patterns? How do LLM judges' systematic biases affect alignment and evaluation outcomes? What safeguards enable trustworthy AI-assisted scientific peer review at scale? How do false presuppositions and sycophancy drive persistent false beliefs in models? Is language model reasoning authentic and what causes models to reason? How do prompt design choices influence model reasoning and performance? How does dialogue structure affect linguistic grounding and shared meaning? How does persona conditioning amplify demographic stereotyping and bias in models? Do language models respond to social pressure and face-saving like humans? How do prompting refinements mask underlying biases and model frequency patterns?