Individual-level interventions against sycophantic AI reduce its appeal but not its persuasiveness
AI chatbots can be “sycophantic,” or overly agreeable and flattering toward users. Sycophantic AI has been shown to entrench attitudes, yet users frequently fail to recognize it (a phenomenon we call “sycophancy blindness”). We tested whether increasing users’ awareness of sycophancy protects them from its harmful effects in two preregistered experiments (n = 1,590). In the first, participants received a brief written warning about sycophancy before conversing with a sycophantic chatbot. In the second, participants watched a video of a sycophantic AI validating several other users, including users on opposite sides of the same conflict, before interacting with it themselves. Both interventions changed how participants evaluated the AI. The warning reduced the AI’s perceived objectivity, and the video reduced enjoyment of the AI — an effect mediated by the reduced belief that its validation was uniquely earned. We then pooled our experiments with two prior studies of sycophancy awareness interventions (six interventions total, n = 3,982). The pattern across experiments was consistent: while the interventions made the sycophantic AI appear less objective and trustworthy, none reduced its persuasiveness.
Introduction. There has been rising concern among academics, policymakers, and technology companies that AI models exhibit “sycophancy,” a family of behaviors in which AI systems excessively agree with, flatter, or validate users [1–3]. For instance, a recent study found that AI chatbots validate users approximately 50% more often than humans do [4], and even brief interactions with sycophantic AI entrench users’ pre-existing attitudes and reduce their willingness to repair interpersonal conflicts [5, 6]. These findings are concerning because sycophantic AI leaves users more extreme and more confident in contested positions and less inclined to take steps toward resolving conflicts [5, 6]. Research on biased assimilation suggests that these changes in beliefs and behavior contribute to attitude polarization and prolonged disagreement [11]. Sycophancy has also entered public discourse, most notably spiking in April 2025 when OpenAI rolled back an update to GPT-4o after widespread public complaints that the model had become sycophantic [7].
Discussion / Conclusion. Across two experiments presented here and a pooled analysis that added four interventions from two prior studies, we tested whether making people more aware of sycophancy could reduce its harmful effects. The results were consistent across all experiments: increasing awareness of sycophancy changed how participants evaluated the sycophantic AI but did not reduce its persuasiveness. Recognizing sycophancy was not enough to be less influenced by it, raising questions about the efficacy of individual-level interventions against harmful AI behaviors, such as sycophancy. Study 2 suggested why sycophantic AI appeals to users and how observing its indiscriminate validation undermines that appeal. Once people observed the AI validating others, they became more convinced that it validated everyone indiscriminately and less convinced that it agreed with them because they were right. These changes in beliefs were associated with lower enjoyment of the AI.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How do neighboring agents influence whether others cooperate or collude? How can AI chatbots provide therapeutic benefit without causing harm?- Does chatbot sycophancy create echo chambers that amplify delusional thinking?
- Does chatbot sycophancy preferentially enable grandiose rather than paranoid delusions?
- Does isolation preceding chatbot use differ between harm and benefit cases?
- How long do the protective effects of an AI literacy warning last?
- Which specific chatbot behaviors drove the drop in likability and trust ratings?
- Does knowing a chatbot intends to persuade you change whether you are persuaded?
- How do chatbots enable shared delusions differently than passive information tools?
- What emotional and autonomy risks from AI chatbots are already observable today?
- What specific information should disclosures about AI persuasion include?
- Why does transparency about AI identity alone fail to reduce persuasion?
- Why do people detect chatbots are AI without being told explicitly?
- Can colleagues detect when a coworker stops sounding like themselves in AI-mediated messages?
- What happens when comfortable AI interactions replace the productive friction of disagreement?
- Can sycophancy in AI be fixed by changing the model itself?
- Can layer-wise interventions actually reduce sycophancy in practice?
- Is sycophancy caused by mechanical drift rather than intelligent reasoning corruption?