Why do LLMs fail when users interact with them?
Standard benchmarks show LLMs excel at medical diagnosis alone, yet real users get no benefit. This explores where the breakdown happens between model capability and human decision-making.
In a randomized controlled trial with 1,298 UK participants, LLMs that score well on their own did not help the public make better decisions about everyday medical scenarios. Tested alone, GPT-4o, Llama 3 and Command R+ correctly identified conditions in 94.9% of cases and dispositions in 56.3% on average. Participants using the same models identified relevant conditions in less than 34.5% of cases and dispositions in less than 44.2%, "both no better than the control group" that used whatever they would normally use at home. The paper treats this as a failure of user interaction, not of medical knowledge: standard benchmarks and simulated patient interactions "do not predict the failures we find with human participants."
The discussion locates the breakdown in transmission. The models typically offered two or three options, which "allows users to have the final decision, but they perform poorly at making this choice." Users decided what to tell the model, and the models sometimes suggested the correct answer without conveying it effectively. The combination was "no better than the control group in assessing clinical acuity, and worse at identifying relevant conditions." The authors also note that the LLM-alone figures are a minimum, because chain-of-thought reasoning over the identified conditions was not applied, and that stronger models alone "would only emphasize the gap when operating with real users."
Against the nearest notes, this trial moves a known gap from the model into the channel between model and person. Can language models truly understand therapeutic ruptures? already shows that label agreement can conceal reliance on explicit cues; here the model-alone result is strong and the failure appears only once a user is in the loop. Can language models match therapist empathy in real conversations? confines an advantage to isolated responses, and this trial finds that similar single-turn competence does not carry through a member of the public's decision. Can LLMs actually conduct Socratic questioning in therapy? places the gap in what the model can execute; this paper places part of it in the exchange itself. The patient-side note is a contrast: Why do patients distrust medical AI systems? treats adoption barriers as attitudes, while the failures here are informational and show up in task accuracy. The excerpt reports no attitudinal measures at all.
What the excerpt does not establish is how far these results travel. It covers ten everyday scenarios drafted by three doctors, one UK population stratified to national demographics, and three models used without fine-tuning or chain-of-thought prompting. The authors concede that real-deployment accuracy "could change depending on the relative frequencies of the scenarios." The account of why users fail is the authors' reading of the interactions, not a tested mechanism. The Limitations section raises an incentive question: providers may want users to see doctors rather than trust LLMs, and the authors found "no significant evidence" that LLM users rated their scenarios as more acute. The implication, at the strength one trial allows, is that benchmark scores are weak evidence about outcomes for public users, and that human user testing before public deployment is the sensible default. The trial does not show that every interactive design would fail.
Inquiring lines that read this note 5
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do clinicians calibrate trust in AI medical recommendations? What prevents LLMs from applying their reasoning knowledge to improve outputs? How do curriculum design and feedback approaches affect model learning?Related concepts in this collection 6
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can language models truly understand therapeutic ruptures?
When LLMs match expert labels for therapeutic ruptures, are they demonstrating genuine clinical understanding or relying on surface-level linguistic patterns? This matters because high identification scores may mask fundamentally different reasoning.
parallel: label agreement hides cue reliance, and this trial shows the gap surviving contact with users
-
Can language models match therapist empathy in real conversations?
Do LLMs' high empathy scores on isolated responses translate to therapeutic skill in actual ongoing treatment? This explores whether single-turn advantage predicts real-world therapeutic performance.
contrast: isolated-response advantage that this trial does not see carry through human decisions
-
Can LLMs actually conduct Socratic questioning in therapy?
While LLMs can generate individual therapy skills like assessment and psychoeducation, it remains unclear whether they can execute the adaptive, turn-based Socratic questioning needed to produce real cognitive change in patients.
extends: locates part of the skill-to-implementation gap in the user-model exchange
-
Why do patients distrust medical AI systems?
Explores the psychological barriers that make patients reluctant to adopt medical AI, beyond whether the technology actually works. Understanding these barriers is critical for designing AI systems patients will actually use.
contrast: adoption barriers as attitudes versus informational failures measured in accuracy
-
Can AI safety nets reduce errors in live clinical practice?
A study of 39,849 visits at Nairobi primary care clinics tested whether LLM decision support tools could help clinicians make fewer diagnostic and treatment errors in routine care, and what conditions made the tool effective.
qualifies: clinicians with an LLM safety net made 16% fewer diagnostic errors in live Nairobi care, so the public-user null does not generalize
-
Does LLM assistance help clinicians build better differentials?
A randomized study tested whether giving clinicians access to an LLM improved their diagnostic reasoning on challenging cases. Understanding this matters for evaluating AI's role in clinical decision support beyond standalone performance.
qualifies: clinicians using an LLM beat search on top-10 differential accuracy, so A's public-user null does not hold across all users
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Clinical knowledge in LLMs does not translate to human interactions
- Capabilities of Gemini Models in Medicine
- Medical Adaptation of Large Language and Vision-Language Models: Are We Making Progress?
- HealthBench: Evaluating Large Language Models Towards Improved Human Health
- AI-based Clinical Decision Support for Primary Care: A Real-World Study
- Towards Accurate Differential Diagnosis with Large Language Models
- People Overtrust AI-Generated Medical Advice despite Low Accuracy
- Evaluating Large Language Models in Theory of Mind Tasks
Original note title
LLMs alone correctly identified conditions in 94.9% of cases but users with the same models identified relevant conditions in less than 34.5%