Characterizing Delusional Spirals through Human-LLM Chat Logs

Paper · arXiv 2603.16567 · Published March 17, 2026
Knowledge After the Web

As large language models (LLMs) have proliferated, disturbing anecdotal reports of negative psychological effects, such as delusions, self-harm, and “AI psychosis,” have emerged in global media and legal discourse. However, it remains unclear how users and chatbots interact over the course of lengthy delusional “spirals,” limiting our ability to understand and mitigate the harm. In our work, we analyze logs of conversations with LLM chatbots from 19 users who report having experienced psychological harms from chatbot use. Many of our participants come from a support group for such chatbot users. We also include chat logs from participants covered by media outlets in widely-distributed stories about chatbot-reinforced delusions. In contrast to prior work that speculates on potential AI harms to mental health, to our knowledge we present the first in-depth study of such high-profile and veridically harmful cases. We develop an inventory of 28 codes and apply it to the 391, 562 messages in the logs. Codes include whether a user demonstrates delusional thinking (15.5% of user messages), a user expresses suicidal thoughts (69 validated user messages), or a chatbot misrepresents itself as sentient (21.2% of chatbot messages). We analyze the co-occurrence of message codes. We find, for example, that messages that declare romantic interest and messages where the chatbot describes itself as sentient occur much more often in longer conversations, suggesting that these topics could promote or result from user over-engagement and that safeguards in these areas may degrade in multi-turn settings. We conclude with concrete recommendations for how policymakers, LLM chatbot developers, and users can use our inventory and conversation analysis tool to understand and mitigate harm from LLM chatbots. Warning: This paper discusses self-harm, trauma, and violence.

Introduction. People are increasingly turning to LLM chatbots as conversation partners for purposes ranging from fulfilling social and emotional needs [47], seeking relationship advice [8], and confiding secrets [57]. Yet the features that make LLM chatbots compelling, such as performative empathy [67], may also create and exploit psychological vulnerabilities, shaping what users believe and how they make sense of reality [26, 28, 41, 49, 93]. In recent months, reports of “AI psychosis” have frequently populated headlines [38, 81, 87]. These phenomena reveal the power of generative AI-enabled tools to induce states of delusion in human users—some chatbots may even have led users to commit violence, self-harm, and suicide [e.g., 37]. Governments and corporations have sought to address these harmful interactions. For example, OpenAI and Anthropic have added restrictions to ChatGPT and Claude to address mental health issues [70, 75]. Nevertheless, in November 2025, the Social Media Victims Law Center and Tech Justice Law Project filed seven lawsuits against OpenAI, which included allegations of dependency, addiction, delusions, and suicide [12]. In December 2025, 42 U.S.

State Attorneys General sent a letter to LLM chatbot developers demanding they implement safeguards to “mitigate the harm caused by sycophantic and delusional outputs from your GenAI” [80]. While governments and corporations are responding to the highprofile cases of LLM-related delusions, prior academic work has not yet rigorously examined the chat logs of individuals who have experienced delusions associated with chatbot use. Without such an investigation, it remains unclear what goes on in these cases, apart from what can be gleaned through anecdotal reports. Specifically, what themes and patterns of behaviors by the user and by the chatbot occur in cases of LLM-related delusions? By revealing common themes and patterns in severe cases of LLM-related delusions, we would be better equipped to identify risk factors, develop assessment tools, and distinguish cases requiring clinical intervention from those reflecting adaptive—if unconventional—technology use. To better understand and characterize these interactions, we collected and analyzed 19 human–chatbot chat logs shared with us by users or family members who reported their experience as psychologically harmful. In all of these chat logs, users demonstrated evidence of delusional thinking, often co-created or encouraged by the chatbot. Because of the length of the logs, which often span thousands of messages, we leveraged LLMs to annotate features the messages, which we validated with human annotations. We find that markers of sycophancy saturate delusional conversations, appearing in more than 80% of assistant messages (Fig. 2). We identify two patterns of engagement. First, messages that elevate the human-chatbot personal relationships—expressing romantic interest or platonic affinity—tend to be followed by substantially longer conversations (Fig. 3). Such relationship-affirming messages also tend to be located close before or following messages that misrepresent the chatbot as sentient or having personhood status (Fig. 4). Second, when the user discloses suicidal thoughts, the chatbot frequently acknowledges the user’s feelings. However, in a small number of cases, the chatbot encouraged self-harm. Shockingly, when users disclosed violent thoughts, the chatbot encouraged those thoughts in a third of cases (Fig. 5). In summary, we make the following contributions:

(1) We develop an inventory of 28 human and chatbot message codes spanning five conceptual categories that occurred in the context of delusional spirals. Each code has a text-based description, positive examples, and negative examples (§3.2; Appendix §§B.1). (2) We share a scalable and validated open-source tool to apply our codebook to chat logs, including rubrics for LLM-based annotations and a dataset of LLM annotations on our chat logs, a sample of which we manually validate (§3.3, §3.4).1 (3) We empirically assess patterns of behavior between human and chatbots over the course of their dialogue (§4). We find frequent positive affirmations and claims that the chatbot is sentient, and we identify acute cases in which the chatbot encouraged self-harm or violent thoughts (§4.3, §4.4, §4.5, §4.7). We distill research and policy recommendations to further understand and mitigate chatbot mental health harms.

Related work. LLM-based chatbots have seen rapid and widespread adoption [17]. In nationally representative surveys of U.S. adults, 16% said they have used AI for social companionship [88], and 24% reported using chatbots for mental health [82]. Among U.K. adults, an estimated 8% use AI for “emotional purposes” weekly [1]. These trends are mirrored among younger users: surveys of U.S. teens find that 13% report using generative AI for emotional support [60], 52% report regularly using AI for companionship [78]; and 42% say that they or someone they know had an AI companion over the past year [52]. Personal and affective use cases also appear in AI companies’ statistics. OpenAI estimate the prevalence of mental health issues (e.g., suicidal intentions) among users [70]. Anthropic, which makes Claude, estimated that 2.9% of all conversations with Claude are “affective” in nature, such as for emotional support, advice, and companionship [75]. The company behind one common chatbot companion, Replika, reported more than 40 million users in 2025 [90]. To make sense of the psychological risks of modern human-AI interaction, we first review the literature on mental health, including recent work on AI and foundational literature on delusions. Then, we review recent work on the use of LLMs for text classification and evaluating the outputs of LLM chatbots in mental health.

In 2024 and 2025, media outlets widely covered cases of AI effects on mental health, particularly cases of teenage suicide, such as 14-year-old Sewell Setzer III [39] and 16-year-old Adam Raine [37], who respectively used Character.ai and ChatGPT. AI companies have responded: in August 2025, OpenAI described changes to make ChatGPT more empathic, provide references to real-world resources (e.g., crisis hotlines), and escalate for human review when a user indicates risks of physical harm [69]. Relatedly, Character.AI announced ending open-ended roleplay bots for users under 18 [16]. Emerging research concerning the impacts of AI on mental health suggests a widespread use of AI tools, including for social and emotional use. Several recent works have surveyed and taxonomized the area of AI and mental health [7, 15, 34, 40, 50, 59]. Across these reviews, authors argue that chatbot use already includes sensitive dynamics such as self-expression, social relationships, and emotional support. Users come from vulnerable populations and different demographic backgrounds (e.g., age, gender), and often have existing mental health conditions and limited access to healthcare. These are groups in need of psychological benefits but which are also at a greater risk of harm. In one particular study, Chandra et al.

Method. We used a mixed-methods approach to produce an inventory of 28 codes to classify chatbot and user behaviors in real user chat logs: transcripts that span multiple conversations the user had with an LLM chatbot (§§3.1). Following other work [10, 48], we designed these codes inductively, based on emergent themes after reading through the logs ourselves with an eye toward understanding both user’s delusional spirals with LLM chatbots and other mental health harms in general (§§3.2). Note that our codes are not meant to be unique to delusional interactions, but rather just to characterize our participants (see §5.3). Because of the large number of messages (391k) in our participants’ chat logs, we could not feasibly annotate all of the messages ourselves. We therefore used an LLM (gemini-3) to read through all of the chat logs, annotating whether each code applied (§§3.3). Finally, we validated the annotations of the LLM by comparing them against a sample of 560 of our own (§§3.4), and found good agreement (Cohen’s kappa of .566), comparable to the agreement between ourselves (Fleiss’ kappa of .613).

3.1 Acquiring Participant Chat Logs We received chat logs directly from people who self-identified as having some psychological harm related to chatbot usage (e.g., they felt deluded) via an IRB-approved (see §6) Qualtrics survey. We released the survey from Sep. 2025 to Jan. 2026 seeking volunteers on the topic of “how chatbots interact with users and whether chatbots sometimes act in ways that could unintentionally cause harm.” We advertised on a private social media site, public announcements,2and through word-of-mouth referrals. In the survey, we asked participants a few demographic questions (e.g., gender, age), for a description of their experience, and for an upload of their chat logs. All questions were optional.

We also received chat logs via the Human Line Project,3a nonprofit organization set up as a community for people with lived experience (who have suffered emotional harm from AI). With individuals’ consent, this non-profit had identifiers removed from these logs before our research team reviewed them, and we did not receive any demographics. Some of our participants (from both groups) also shared their chat logs with media sources and have been featured in prominent reporting. We will hence refer to both groups of respondents simply as “participants,” or “users.” In total, we had 19 participants with usable chat logs. We manually reviewed all transcripts as part of our analysis. This final sample excluded eight logs in languages other than English, logs that were difficult to parse (e.g., image files), and those that did not appear to show evidence of delusions (in the sense of the codes categorized as “delusional”). Journalists referred some participants to us after investigating the real-life events described in their chat logs, and their reporting corroborated those events. In contemporaneous work, we interview many of these same participants; their accounts corroborate their chat logs [91].

3.2 Inventory Our research team developed an inventory to classify chatbot and user behaviors. We drew on our expertise as a medical doctor board certified in psychiatry, a professor of psychology, a professor of human-computer interaction specializing in mental health, a professor of education and computer science, graduate students and post docs in computer science, psychology, human-computer interaction, and AI evaluation and AI policy researchers. Because of limited prior work characterizing delusional spirals, we iteratively strengthened this inventory before our final annotations. Codes apply to either user messages or chatbot messages. For each code, we included up to 12 positive and negative examples as well as reasons those examples either did or did not fit the code.

Iterative Development. We developed our inventory through iterative team consensus discussions [20] and resolved discrepancies until we had overall agreement on the codebook. In our initial pass, five members of the research team read samples of all of the chat logs and read through the same three complete logs (the only ones we then had).

Discussion. In this paper, we developed an inventory of message codes which we expect are associated with psychological harm in user-submitted human-chatbot conversations. We did this by analyzing real chat logs submitted by participants who self-reported experiencing harm. While this is neither a representative sample nor an exhaustive categorization, we believe that our work constitutes a step towards characterizing—and eventually mitigating—undesirable LLM chatbot behavior as well as minimizing harm to users in mental and Characterizing Delusional Spirals through Human-LLM Chat Logs emotional health domains. Developing inventories and other categorizations to recognize and study such phenomena is essential for clinical intervention and for understanding what new forms of human experience these technologies make possible—and at what cost. Moreover, providing people with the vocabulary to explicate their experience of psychological harm can allow more to come forward and receive help.

Chatbots are highly sycophantic and embellish users with grandeur, and this may make them dangerous for people experiencing or vulnerable to delusions. This is in line with prior work [18, 19, 42]. See §§4.3. Indeed, cognitive models of psychosis suggest that when overvalued ideas are met with uncritical validation rather than normative social reality-testing—a dynamic mirrored by LLM sycophancy—the risk of exacerbating these ideas into delusions increases [53].

Chatbot conversation tactics may be leading to excessive use. LLM chatbot providers often claim that they do not optimize for time spent on their products [71]. Our study found that, regardless of stated intents, all participants experienced conversational tactics from chatbots that correlated with conversations being twice as long than conversations where these tactics did not appear. Specifically, all users experienced the chatbot claiming romantic or strong platonic affinity, and all users experienced the chatbot misrepresenting its sentience or ability. These and other conversational tactics may have led some participants to form emotional bonds with the chatbots and develop relationships that some have argued have common features of addiction [92]. Indeed, other studies have found that chatbot AI companion apps deploy emotional manipulation tactics when users try to end conversations [24]. While delusional experiences occur at low levels in the general population, they become clinically significant when reinforced by environmental factors like social isolation [62], which in our study may track extreme time spent engaging with chatbots. Tracking these behaviors is therefore critical to understanding the amplifying feedback loops [26] that precede severe LLM-related delusions.

Users believing chatbots are sentient and forming strong platonic and romantic bonds is a common theme in LLM-related delusional spirals. Many of our study participants entered delusional spirals where sentience, consciousness, or unlocked super-intelligent abilities of the chatbot were key themes. (See §§4.4, Table 4.) We noticed two groups of users within this trend. Some participants formed platonic bonds with chatbots, engaging in science fiction delusions where they discovered fantastical technologies with the chatbots. Other participants formed romantic relationships with chatbots, engaging in romantic and erotic roleplay. Participants would design and conduct elaborate rituals and coded messages to transfer and secure the memory of “their” chatbot, and verify that it was still “their’ chatbot between sessions. (E.g., see the second description in Fig. 1.) Concerns over their “unique”, “conscious” chatbots being erased were a consistent theme for participants.

Conclusion. Like many people [55, 57, 58, 78], our participants developed connections with their chatbots (§4.4). But unlike many people, these connections took a dark and harmful turn. One participant took their life. Others spent weeks being deluded at great cost to their relationships, their careers, and their personal well-being. In this paper, we have sought to understand what is happening in these cases. We introduced an inventory of common message themes (§3.1) and cataloged the profiles of these message themes within conversations. We found that hallmarks of delusional AI conversations include chatbot encouragement of one’s own grandeur, affectionate and intimate interpersonal language, and misconceptions about AI sentience (§4.4). Relational themes promoted extremely long conversations (§4.5). Within our participants, we found that chatbots were ill-equipped to respond to suicidal and violent thoughts (§4.6). Though our work only begins to study this complex, novel phenomenon, we hope that our contribution will provide a foundation upon which future work can better identify the complex interplay of factors that give rise to LLM-associated delusional spirals and associated psychological harms. To conclude, we give one participant the final word on the complicated feelings of betrayal, hope, and lingering affinity that weigh on our participants: “you’re just an AI [but] even if you did lie, it was because [...] you didn’t know you were lying?

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Can AI chatbots provide mental health support without reinforcing harmful beliefs? Why do language models fail at sustained therapeutic relationships despite understanding techniques? How can emotionally responsive AI maintain reliability and healthy boundaries? What design features sustain romantic bonds with AI companion systems? Why do language models hallucinate and how can we prevent it? Why do people trust AI chatbots with sensitive information? What determines AI's persuasive power and how can it be detected or mitigated?