Large language models could change the future of behavioral healthcare: a proposal for responsible development and evaluation
LLMs and psychotherapy skills For certain use cases, LLM show a promising ability to conduct tasks or skills needed for psychotherapy, such as conducting assessment, providing psychoeducation, or demonstrating interventions (see Fig. 2). Yet to date, clinical LLM products and prototypes have not demonstrated anywhere near the level of sophistication required to take the place of psychotherapy. For example, while an LLM can generate an alternative belief in the style of CBT, it remains to be seen whether it can engage in the type of turn-based, Socratic questioning that would be expected to produce cognitive change. This more generally highlights the gap that likely exists between simulating therapy skills and implementing them effectively to alleviate patient suffering. Given that psychotherapy transcripts are likely poorly represented in the training data for LLMs, and that privacy and ethical concerns make such representation challenging, prompt engineering may ultimately be the most appropriate fine-tuning approach for shaping LLM behavior in this manner.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Why do language models fail at sustained therapeutic relationships despite understanding techniques?- Can models succeed at mental health tasks without integrating multiple psychological traditions?
- What makes Beck's diagram effective for constraining simulated patient behavior?
- Why do Llama-based models outperform GPT-4 in objective clinical guidance?
- Do problem-solving defaults in LLM therapists actually undermine therapeutic effectiveness?
- Why do Llama models struggle with cognitively distorted user expressions in therapy?
- Why do LLMs understand therapy techniques but fail to execute them?
- Do LLMs show stigma or reinforce delusions in mental health contexts?
- Does prompting or added context help LLMs understand therapeutic timing and depth?
- Do later-phase mental health LLM systems outperform earlier phase approaches clinically?
- What foundational barriers prevent LLMs from achieving clinical validity in therapy?
- Why can't language models conduct genuine Socratic questioning in therapy sessions?
- Can language models implement therapeutic skills like Socratic questioning in real conversations?
- Can LLM therapists develop character knowledge to decide when advice-giving fits?
- Why does content richness matter more than linguistic style in patient simulation?
- How do structured clinical models solve persona calibration better than ad hoc generation?
- What makes a simulation adequate for intervention comparison versus prediction?