INQUIRING LINE

AI scored higher than trainee therapists on single replies, but that test never checked whether the edge holds over a long conversation.

Why do single-turn LLM responses outperform humans while ongoing relationships show limits?

This explores why LLMs can beat people when judged on one reply at a time, like a single empathetic answer to a therapy client, yet seem to fall short in long conversations and ongoing relationships, and whether the corpus shows that gap or only suspects it.


This explores why LLMs look better than humans when you grade one reply at a time, but look weaker once the job becomes an ongoing conversation or relationship. One fact matters before anything else: the clearest head-to-head result in the corpus never tested the relationship side. In that study, six LLMs scored higher than trainee therapists on empathy, validation and clinical knowledge in isolated responses. The authors say the advantage holds only for single-turn evaluation, and multi-turn therapy and real outcomes were not measured Can language models match therapist empathy in real conversations?. So the 'limits' are not a measured failure in therapy. They are a gap in the evidence. To see why that gap is worrying, you have to look at what other parts of the corpus show about long conversations.

The strongest clue comes from task research, not therapy. Models score about 90% when an instruction arrives in one complete message, but only about 65% when the same information comes out bit by bit over a natural conversation. They lock onto an early guess and can't back out of it. The authors trace this to RLHF training, which rewards answering helpfully over asking clarifying questions Why do AI assistants get worse at longer conversations?. Therapy is the most extreme version of 'information arrives gradually': a client rarely states the real problem in the first session. The same helpfulness pull shows up directly in clinical-style tests. When users share emotions, LLM therapists jump to problem-solving, a known sign of low-quality therapy, even though they also reflect on clients' strengths more than poor human therapists do Do LLM therapists respond to emotions like low-quality human therapists?. A single reply can look excellent while carrying a habit that would hurt over months.

A second explanation is about what a single reply actually is. An LLM doesn't commit to one stable character. It holds a spread of possible personas that narrows as the conversation goes on, and each reply is a draw from that spread Does an LLM commit to a single character or maintain many?. A one-shot evaluation grades one draw. Setting temperature to zero doesn't fix that: you get the same output every time, but it is still just one sample, so consistency is not the same as reliability Does setting temperature to zero actually make LLM outputs reliable?. A relationship, by contrast, asks the model to *be* someone steady over time, which is a different and harder test.

A third explanation is that single-turn tests leave out the hard part of social life: not knowing what the other person knows. LLMs look socially skilled when one model plays every side of a conversation, but they fail in consistent ways when each party holds private information. Their apparent competence depended on skipping the work of building shared understanding Why do LLMs fail when simulating agents with private information?. A therapist, friend or advisor does that work all the time. Grading one isolated reply, written with full context handed over, quietly removes it.

Persuasion research suggests a broader lesson: whether an LLM beats humans depends on the setup, not on the speaker being a machine. Pooled across 17,000+ participants, LLMs and humans are equally persuasive Are language models actually more persuasive than humans?. Individual wins depend on the model and on whether it argues for truth or falsehood Do large language models persuade better than humans?. One more finding is easy to miss: LLMs try to persuade in nearly every exchange, using logic and numbers that sound objective Do LLMs persuade users more often than humans do?. Taken one reply at a time, that reads as clear and competent. Repeated across a long relationship, it could amount to a constant, unearned pull on what you believe. The pattern across the corpus is that the behaviors that win a single turn (confident, helpful, polished) are often the ones that cause trouble over many turns.


Sources 9 notes

Can language models match therapist empathy in real conversations?

Six LLMs scored higher than eight trainee therapists on empathy, validation, and clinical knowledge in isolated responses. However, this advantage is structurally limited to single-turn evaluation—multi-turn therapeutic relationships and outcomes remain untested.

Why do AI assistants get worse at longer conversations?

LLMs perform at 90% accuracy with single-message instructions but drop to 65% across natural conversation. Models lock into early guesses when information arrives gradually and cannot course-correct, a behavior induced by RLHF training that rewards helpfulness over clarification.

Do LLM therapists respond to emotions like low-quality human therapists?

Using the BOLT framework, researchers found LLMs offer solution-focused advice during emotional disclosure—a hallmark of low-quality therapy—yet also reflect more on client needs and strengths than typical poor human therapy, creating an unusual hybrid profile likely driven by RLHF's helpfulness bias.

Does an LLM commit to a single character or maintain many?

Research shows LLMs don't commit to a single character but instead maintain a probability distribution over many consistent simulacra. Each response samples from this distribution, explaining why regenerations can yield different personalities while remaining consistent with prior context.

Does setting temperature to zero actually make LLM outputs reliable?

Fixed seeds and zero temperature replicate the same output repeatedly, but that output remains one draw from the model's probability distribution. McDonald's omega testing across 100 repetitions reveals that consistency does not equal reliability.

Show all 9 sources
Why do LLMs fail when simulating agents with private information?

Research shows LLMs perform well when one model controls all interlocutors but fail systematically when agents possess private information. This reveals that apparent social competence relies on grounding work that models skip in omniscient settings.

Are language models actually more persuasive than humans?

A meta-analysis of 7 studies with 17,422 participants found no detectable difference in persuasive effectiveness between LLMs and humans (Hedges' g = 0.02). Persuasiveness appears conditional on context rather than speaker category.

Do large language models persuade better than humans?

Claude beats incentivized humans at both truthful and deceptive persuasion, while DeepSeek only beats them when arguing for falsehoods. The persuasion mechanism appears content-independent, suggesting model family itself acts as a contextual moderator.

Do LLMs persuade users more often than humans do?

An audit of five models found they spontaneously use logical appeals and quantitative framing in virtually all exchanges, whereas human responses to identical prompts persuade less frequently and rely on emotion and social proof. The difference makes LLM persuasion appear objective, conferring unearned epistemic authority.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.