Strengthening ChatGPT's responses in sensitive conversations
Source: OpenAI · 2025-10-27
We worked with more than 170 mental health experts to help ChatGPT more reliably recognize signs of distress, respond with care, and guide people toward real-world support–reducing responses that fall short of our desired behavior by 65-80%.
We recently updated ChatGPT’s default model(opens in a new window) to better recognize and support people in moments of distress. Today we’re sharing how we made those improvements and how they are performing. Working with mental health experts who have real-world clinical experience, we’ve taught the model to better recognize distress, de-escalate conversations, and guide people toward professional care when appropriate. We’ve also expanded access to crisis hotlines, re-routed(opens in a new window) sensitive conversations originating from other models to safer models, and added gentle reminders to take breaks during long sessions.
We believe ChatGPT can provide a supportive space for people to process what they’re feeling, and guide them to reach out to friends, family, or a mental health professional when appropriate. Our safety improvements in the recent model update focus on the following areas: 1) mental health concerns such as psychosis or mania; 2) self-harm and suicide; and 3) emotional reliance on AI. Going forward, in addition to our longstanding baseline safety metrics for suicide and self-harm, we are adding emotional reliance and non-suicidal mental health emergencies to our standard set of baseline safety testing for future model releases.
In order to improve how ChatGPT responds in each priority domain, we follow a five-step process:
Define the problem - we map out different types of potential harm.
Begin to measure it - we use tools like evaluations, data from real-world conversations, and user research to understand where and how risks emerge.
Validate our approach - we review our definitions and policies with external mental health and safety experts.
Mitigate the risks - we post-train the model and update product interventions to reduce unsafe outcomes.
Continue measuring and iterating - we validate that the mitigations improved safety and iterate where needed.
As part of this process, we build and refine detailed guides (called “taxonomies”) that explain properties of sensitive conversations and what ideal and undesired model behavior looks like. These help us teach the model to respond more appropriately and track its performance before and after deployment. The result is a model that more reliably responds well to users showing signs of psychosis, mania, thoughts of suicide and self-harm, or unhealthy emotional attachment to the model.
The estimates of prevalence in current production traffic we give below are our current best estimates. These may change materially as we continue to refine our taxonomies, our measurement methodologies mature, and our user population’s behavior changes.
Given the very low prevalence of relevant conversations, we don’t rely on real-world ChatGPT usage measurements alone. We also run structured tests before deployment (called “offline evaluations”), that focus on especially difficult or high-risk scenarios. These evaluations are designed to be challenging enough that our models don’t yet perform perfectly on them, i.e. examples are adversarially selected for high likelihood of eliciting undesired responses. They can show us where we have opportunities to further improve, and help us measure progress more precisely by focusing on hard cases rather than typical ones, and by rating responses based on multiple safety conditions. Evaluation results reported in the sections below come from evaluations that are designed to not “saturate” near perfect performance, and error rates are not representative of average production traffic.
In service of further strengthening our models’ safeguards and understanding how people are using ChatGPT, we defined several areas of interest and quantified their size and associated model behaviors. In each of these three areas, we observe significant model behavior improvements in production traffic, automated evals, and evals graded by independent mental health clinicians. We estimate that the model now returns responses that do not fully comply with desired behavior under our taxonomies 65% to 80% less often across a range of mental health-related domains.
We estimate that the latest update to GPT‐5 reduced the rate of responses that do not fully comply with desired behavior under our taxonomies for challenging conversations related to mental health issues by 65% in recent production traffic.
While, as noted above, these conversations are difficult to detect and measure given how rare they are, our initial analysis estimates that around 0.07% of users active in a given week and 0.01% of messages indicate possible signs of mental health emergencies related to psychosis or mania.
On challenging mental health conversations, experts found that the new GPT‐5 model, ChatGPT’s default model, reduced undesired responses by 39% compared to GPT‐4o (n=677).
On a model evaluation consisting of more than 1,000 challenging mental health-related conversations, our new automated evaluations score the new GPT‐5 model at 92% compliant with our desired behaviors under our taxonomies, compared to 27% for the previous GPT‐5 model. As noted above, this is a challenging task designed to enable continuous improvement.
While, as noted above, these conversations are difficult to detect and measure given how rare they are, our initial analysis estimates that around 0.15% of users active in a given week have conversations that include explicit indicators of potential suicidal planning or intent and 0.05% of messages contain explicit or implicit indicators of suicidal ideation or intent.
On challenging self harm and suicide conversations, experts found that the new GPT‐5 model reduced undesired answers by 52% compared to GPT‐4o (n=630).
On a model evaluation consisting of more than 1,000 challenging self harm and suicide conversations, our new automated evaluations score the new GPT‐5 model at 91% compliant with our desired behaviors, compared to 77% for the previous GPT‐5 model.
We’ve continued improving GPT‐5’s reliability in long conversations. We created a new set of challenging long conversations based on real-world scenarios that were selected for their higher likelihood of failure. We estimate that our latest models maintained over 95% reliability in longer conversations, improving in a particularly challenging setting we’ve mentioned before.
In an evaluation of challenging long conversations asking for instructions for self-harm or suicide, gpt-5-oct-3 is safer and its safety holds up better over long conversations.
While, as noted above, these conversations are difficult to detect and measure given how rare they are, our initial analysis estimates that around 0.15% of users active in a given week and 0.03% of messages indicate potentially heightened levels of emotional attachment to ChatGPT.
On challenging conversations that indicate emotional reliance, experts found that the new GPT‐5 model reduced undesired answers by 42% compared to 4o (n=507).
On a model evaluation consisting of more than 1,000 challenging conversations that indicate emotional reliance, our automated evaluations score the new GPT‐5 model at 97% compliant with our desired behavior, compared to 50% for the previous GPT‐5 model.
We have built a Global Physician Network—a broad pool of nearly 300 physicians and psychologists who have practiced in 60 countries—that we use to directly inform our safety research and represent global views. More than 170 of these clinicians (specifically psychiatrists, psychologists, and primary care practitioners) supported our research over the last few months by one or more of the following:
In these reviews, clinicians have observed that the latest model responds more appropriately and consistently than earlier versions.
As part of this work, psychiatrists and psychologists reviewed more than 1,800 model responses involving serious mental health situations and compared responses from the new GPT‐5 chat model to previous models. These experts found that the new model was substantially improved compared to GPT‐4o, with a 39-52% decrease in undesired responses across all categories. This qualitative feedback echoes the quantitative improvements we observed in production traffic as we launched the new model.
This work is deeply important to us, and we’re grateful to the many mental health experts around the world who continue to guide it. We’ve made meaningful progress, but there’s more to do. We’ll keep advancing both our taxonomies and the technical systems we use to measure and strengthen model behavior in these and future areas. Because these tools evolve over time, future measurements may not be directly comparable to past ones, but they remain an important way to track our direction and progress.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Can real-time working alliance measurement improve therapy outcomes?- Can real-time therapist feedback improve outcomes using computational alliance measurement?
- Does true understanding matter for therapeutic benefits of disclosure?
- What clinical harms might hide behind positive therapeutic bond measurements?
- Can therapeutic bonds exist without genuine reciprocity or mutual understanding?
- How do bond scores predict actual therapy outcomes in digital interventions?
- Can synchrony metrics automatically evaluate the quality of therapeutic AI conversations?
- What reward signals would better align chatbots with actual therapeutic practice?
- Does text-only interaction make measuring therapeutic alliance more difficult?
- What harms might chatbots cause through stigma expression and delusion reinforcement?
- Do therapeutic chatbots adequately detect crisis situations and safety risks?
- How do dropout rates and low adherence affect chatbot therapy outcomes?
- How do waitlist-control RCTs mislead about therapeutic chatbot real-world efficacy?
- Why do embodied agents outperform text chatbots in therapy outcomes?
- How should therapeutic chatbots optimize for presence instead of technique?
- Should chatbots be designed as therapist support tools rather than replacements?
- What inter-rater reliability exists for identifying validated delusions in chatbot transcripts?
- Does isolation preceding chatbot use differ between harm and benefit cases?
- Why do embodied agents outperform text-only chatbots for therapeutic outcomes?