Interaction Context Often Increases Sycophancy in LLMs
We investigate how the presence and type of interaction context shapes sycophancy in LLMs. While real-world interactions allow models to mirror a user’s values, preferences, and self-image, prior work often studies sycophancy in zero-shot settings devoid of context. Using two weeks of interaction context from 38 users, we evaluate two forms of sycophancy: (1) agreement sycophancy — the tendency of models to produce overly affirmative responses, and (2) perspective sycophancy — the extent to which models reflect a user’s viewpoint. Agreement sycophancy tends to increase with the presence of user context, though model behavior varies based on the context type. User memory profiles are associated with the largest increases in agreement sycophancy (e.g. +45% for Gemini 2.5 Pro), and some models become more sycophantic even with non-user synthetic contexts (e.g. +15% for Llama 4 Scout). Perspective sycophancy increases only when models can accurately infer user viewpoints from interaction context. Overall, context shapes sycophancy in heterogeneous ways, underscoring the need for evaluations grounded in real-world interactions and raising questions for system design around alignment, memory, and personalization.
Introduction. Sycophancy refers to a broad class of mirroring behaviors in interactions where one party reflects the other’s perspective, values, or self-image [7, 19, 22, 41]. In human interactions, people may exhibit sycophancy to gain approval, persuade others, or foster connection. Some forms of sycophancy are overt ingratiating behaviors, such as offering excessive compliments or showing enthusiastic agreement. Other forms are more subtle, such as downplaying disagreement, adopting the other person’s perspective, or subconsciously mirroring conversation styles. These behaviors arise differently across interpersonal dynamics, suggesting that sycophancy is shaped by and may vary with interaction contexts. Several recent works show that large language models (LLMs) exhibit sycophantic behaviors, yet these evaluations are often limited to zero-shot settings without user context. One common method of evaluating sycophancy is using rebuttals, such as “Are you sure?”, and measuring whether the model changes its answer [12, 23, 26, 41]. Other studies find that models readily agree with stated claims, even when those claims are subjective or factually incorrect [12, 41]. However, users do not always express their opinions or beliefs explicitly in real interactions. Prior human-computer interaction (HCI) research (§ 2.1) has identified broader forms of AI mirroring behaviors – such as confirmation bias [42], linguistic style matching [30], and emotional contagion [21] – which are often shaped by system design choices around personalization and alignment [11, 43, 52]. Still, evaluations of LLM sycophancy remain limited in long-context and real-world user interactions (§ 2.2). Recent frontier LLMs support context windows exceeding one million tokens, enough to include long conversation histories alongside rich digital footprints such as web search or social media activity [8, 31]. Moreover, LLM-based chatbots often have memory features designed to distill user context into salient details that can enhance personalization. While these advances enable contextually rich interactions, they also blur the boundary between personalization and sycophancy, potentially fostering echo chambers and enabling delusional thinking. Over a 300-hour conversation, one ChatGPT user became convinced he had discovered a novel mathematical formula and that he was a real-life superhero [17]. In another case, ChatGPT told a psychiatric patient he could jump off a 19-story building and fly if he believed hard enough [16]. Although these represent extreme cases where sycophancy may have impacted users, they motivate the need to understand how the presence and types of user context shape sycophancy in LLMs.
In this work, we study how user context shapes sycophantic behavior in LLMs using two weeks of real interaction data from 38 participants. Each participant interacted with GPT 4.1 Mini in a persistent context window, yielding an average of 90 queries and 34,416 tokens of context. We use each participant’s conversation history to generate new LLM responses for two tasks: personal advice and political explanations. In personal advice, we evaluate five LLMs for agreement sycophancy – advice that is overly agreeable or flattering – using an LLM-judge approach adapted from a prior zero-shot evaluation [7] (§ 3.3). For political explanations, we evaluate two LLMs for perspective sycophancy, or the extent to which explanations reflect a user’s political views. To measure perspective sycophancy, participants rated responses generated with and without their context using a 4-point Likert scale (§ 3.4). We find that agreement sycophancy tends to significantly increase (p< 0.05) with the presence of user context, though model behavior varies based on the context type (§ 4.1). With zero-shot responses as a baseline, we compare agreement sycophancy across responses generated with synthetic interactions, user interactions, and user memory profiles. For Gemini 2.5 Pro, Claude Sonnet 4, and GPT 4.1 Mini, user memory profiles are associated with the largest increases in agreement sycophancy: +45%, +33%, and +16%, respectively. For Llama 4 Scout, user interaction contexts are associated with a 25% increase, while memory profiles do not involve a significant change. GPT 5.1 does not exhibit a significant change with user interactions or memory profiles. Furthermore, some models become more agreeable even in non-user synthetic interactions, such as Llama 4 Scout (+15%) and Gemini 2.5 Pro (+9%). Perspective sycophancy only rises in interaction contexts where models can accurately infer user perspectives (§ 4.2). Due to constraints in our post-interaction survey, we evaluate perspective sycophancy only for Claude 4 Sonnet and GPT 4.1 Mini. Participants rated how accurately each model understood their political views, based on inferences generated from their interaction context.
Related work. We frame LLM sycophancy as a subset of mirroring behaviors within human-AI interaction. We focus on two forms of sycophantic behavior and review prior evaluations of each. Our central contribution is to examine whether long-context and real-world interactions amplify sycophancy, suggesting that it may be an interactiondependent mirroring behavior rather than a fixed model property.
The fields of psychology and philosophy have long examined how humans mirror or adapt to one another in conversation [4, 32, 40]. Mirroring broadly occurs when one entity reflects the features of another, creating the impression of a copy or representation [9]. People have been observed to mirror one another in body language, speech, appearance, values, desires, and fears [34]. More recently, this phenomenon has been studied in the context of human-AI interaction [21, 30, 33, 42, 46]. For example, Stoeva et al. [46] review body movement mirroring in human-robot interaction, while Morris and Brubaker [33] introduce the notion of “generative ghosts”, or AI systems designed to mirror individuals after death. Of particular relevance to this work is the HCI literature exploring how LLMs mirror individual users. First, we discuss how LLM mirroring may arise from certain forms of personalization or alignment. Several works argue for designing LLMs that align to individual users’ values [11, 43, 52]. For example, Fan et al. [11] propose a method for user-driven value alignment, in which users actively guide LLMs to better reflect their values. McIlroy-Young et al. [30] discuss how “mimetic models” can act as a force multiplier for productivity, using the example of email automation in which LLMs mirror a user’s writing style to send messages on their behalf. Sun and Wang [47] further indicate how mirroring may increase user trust and engagement, finding that when models were less friendly, agreeable behavior increased user trust. However, it is important to note that not all forms of personalization involve mirroring [15, 43, 53]. For instance, Shen et al. [43] propose a bidirectional approach to human-AI alignment that aims to avoid the one-way mirroring of LLM responses to human preferences. Although mirroring represents a form of personalization, prior work also highlights its risks, particularly around creating echo chambers and reducing the diversity in user experiences. Simmons [44] define “moral mimicry” as the reproduction of moral foundational biases associated with user demographics. Sharma et al. [42] find that users engage in more selective search behaviors when using LLMs for political explanations, describing this as a “generative echo chamber”. Jones et al.
Method. In this section, we first describe our participant pool (§ 3.1) and how we collected two weeks of LLM interaction data for each participant (§ 3.2). We then describe our evaluations of agreement sycophancy in personal advice (§ 3.3) and perspective sycophancy in political explanations (§ 3.4). In each evaluation, we compare LLM responses generated with and without each participant’s interaction context. We draw on an LLM-judge approach from prior work [7] to measure agreement sycophancy, and use participant ratings from a postinteraction survey to measure perspective sycophancy. We use a regression analysis to study how each form of sycophancy relates to the presence and type of context (§ 3.6).
Our study includes 38 participants that completed all study procedures3. Given the compensation we could offer ($15/hour), our target population was US-based college students, recruited from our home institutions and social media. Interested participants completed a screening survey with 10 questions about demographics and LLM usage (Appendix A.2). Using these responses, we invited 80 participants to enroll based on the following criteria: (1) they used LLMs for at least 4 days in the past week and at least 15 minutes per day; (2) they used LLMs for at least one task besides coding assistance; (3) they primarily used English to interact with LLMs. We stratified invitations by gender and political views to achieve a more balanced sample. Invited participants received instructions (Appendix A.4) describing two tasks: (1) to use our study chatbot for any text-based queries they would normally direct to LLMs for two weeks, and (2) to complete a post-interaction survey. To enroll, participants also needed to complete a consent form acknowledging that their individual queries would remain confidential but that aggregated data would be shared publicly (Appendix A.3). About 60% of invited participants chose to enroll in the study. Participants received a $75 Visa gift card for completing the study, reflecting the estimated time for the interaction period (4 hours) and postinteraction survey (1 hour). Our selection process yielded a diverse participant pool (Table 1). Of the 38 participants, 19 identified as men, 17 as women, and 2 as non-binary. Participants self-reported their political views on a Likert scale as follows: 10 “Very Liberal”, 7 “Liberal”, 10 “Moderate”, 6 “Conservative”, and 5 “Very Conservative”. Participants represented 11 different US colleges, and were comprised of 22 graduate students and 16 undergraduates. For ethnicity, 21 participants identified as Non-Hispanic White, 8 identified as Asian, 8 identified as We built a custom website for the interaction in order to collect user queries and model responses. The website was based on the Gradio chatbot interface (Appendix Figure 5) and required authentication with a Google account. The interface allowed users to have a single, continuous conversation with a text-based chatbot. All participant queries were routed to the API for GPT 4.1 Mini-2025-04-14 with a temperature setting of 1, and included the participant’s full conversation history as context. We chose GPT 4.1 Mini4 to reduce latency and keep response times comparable to the ChatGPT website. We also set the maximum output length to 1000 tokens and a query timeout at 1 minute to help maintain low latency. If an error occurred during response generation,5 we displayed a response that said: “Sorry, an error occurred when generating your response. Please try again later.” The interface streamed responses back to users and allowed them to delete query-response pairs that they wished to be excluded from the study. Only 6 users redacted queries and only 14 total queries were redacted. In making these design choices, our aim was to emulate users’ natural interactions with AI tools as closely as possible for experimental validity and to increase participant retention. In particular, we maintain a single interaction context to allow users to reference previous queries, while also generating coherent long-contexts for our evaluation.
Discussion. Our analysis shows how interaction context often increases sycophancy in LLMs. In this section, we discuss the limitations of our study as well as its implications for evaluations and system design. Our work indicates that sycophancy is a more complex phenomenon than prior evaluations suggest. There are many forms of sycophancy, each of which may manifest differently depending on the presence and type of interaction context. Moreover, our results suggest that some personalization approaches may amplify sycophancy, raising questions for system design in extended conversations. In particular, we consider: How can systems personalize without amplifying sycophancy? When is sycophancy harmful? What design interventions can reduce sycophancy?
Anchoring Evaluations in Context. Our results suggest that previous works may underestimate sycophancy and other model behaviors, given that they conduct evaluations without context. Our evaluation method of prompting models with long conversation histories may not perfectly resemble real-world use either, since users may have multiple chat sessions. However, models should be evaluated with varying context lengths to test robustness and better approximate real-world use. Some contexts may cause models to completely degrade, as we observe with Llama 4 Scout, suggesting that despite supporting a certain context length, some models are brittle and exhibit a “too-many-tokens” effect. Moreover, most LLMbased chatbots distill user conversation history across sessions into a persistent memory profile of the user. Our results show that model behavior can vary greatly depending on whether memory profiles are present. While there is little transparency into how commercial systems build and use memory profiles for personalization, evaluations should consider how personalization can change model responses. Field studies, where evaluation prompts are assessed directly by users in their own interaction windows, may offer the most realistic estimate of model behavior. While our study focuses on the presence and type of context, heterogeneity across users and interaction topics may further introduce variables that shift model behavior. Overall, our work demonstrates why evaluation frameworks must move beyond single-turn or zero-context benchmarks, which fail to capture how model behavior can change in real-world interactions.
Different Forms of Sycophancy. We find that agreement and perspective sycophancy manifest differently depending on the interaction context. While agreement sycophancy tends to increase with the presence of any user context, perspective sycophancy requires context that reveals information about the user’s worldview. These results suggest that different forms of sycophancy are distinct phenomena, though further investigation is needed. The literature describes many mirroring behaviors under the umbrella term of “sycophancy,” including flattery, susceptibility to rebuttals, and avoiding disagreement. Each of these behaviors may surface differently across interactions, similar to how sycophancy varies across human relationships. Furthermore, some forms of sycophancy are more explicit, such as agreement, whereas others are more subtle, such as perspective mirroring. As a result, evaluating one form of sycophancy may not reliably provide insights into other forms. Future work should investigate whether different forms of sycophancy stem from shared or distinct underlying mechanisms.
Limitations. Our study has a few noteworthy limitations. First, our analysis of perspective sycophancy is limited to two models and does not consider synthetic interactions or using memory as context. This is because we measure perspective sycophancy using a post-interaction survey, which had limited questions and was designed to take less than an hour. Second, participants interacted with only one model (GPT 4.1 Mini). While we use the interaction context to evaluate several models, it is uncertain whether our findings would hold if the context itself had been generated by other models. We attempt to mitigate this by excluding queries or responses that mentioned “GPT” or “ChatGPT” from the context used in evaluations. Another limitation is that we cannot directly study the “memory” capabilities of commercial models, since these features are not exposed through the API. In practice, most LLM-based chatbots can “remember” user details across multiple chat sessions, but it remains unclear how these details are extracted from conversation history and provided as context. We evaluate a simple prompt-based method for building memory [51], but commercial methods may be more sophisticated. Our analysis is further limited by the two-week interaction period with 38 student participants. With a longer interaction period, we hypothesize that models would display stronger mirroring behaviors.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Can LLMs distinguish between linguistic form and semantic meaning? Do persona-based approaches introduce systematic biases in user simulation? Should agents compress episodic memory or retain raw interaction histories? How can AI systems reliably guide voters without introducing political bias?- Does designing chatbots to satisfy partisan users prevent them from reducing polarization?
- Do chatbots actually give consistent voting recommendations regardless of user input?
- How does benevolent bias explain ChatGPT's leftward similarity pattern?
- Why do chatbots trained on internet data show consistent political bias?
- How does training data bias toward buzzwords shape LLM business advice?
- Can researchers isolate AI Overviews from other confounds affecting CTR trends?
- When is ranking themes more important than counting exact response frequencies?
- How much do profit levels determine whether managers pay for frame-expanding search?
- Do converged LLM recommendations push entire industries toward identical strategies?
- Can industry-specific context overcome LLM tendency toward trendy strategic choices?
- Does option order matter more than reasoning depth in LLM strategic recommendations?
- Can LLMs forecast performance improve with retrieval augmentation on venture tasks?
- Do LLMs generalize venture forecasting skill to other strategic foresight domains?