Can clinicians tell GPT-4 advice apart from expert advice?
This study explores whether trained clinicians can distinguish AI-generated psychological advice from expert advice, and how they rate the quality and empathy of each. The question matters for understanding whether AI might reliably supplement human expertise in mental health settings.
In a blinded comparison, licensed clinicians rated GPT-4 advice as equal or more favorable than expert advice from a Swedish newspaper advice column, on empathy and scientific quality, and they could not reliably tell the two apart. The excerpt reports that "AI-generated advice received equal or more favorable ratings across all measures" over 104 response pairs rated by 43 clinicians (40 psychologists, three psychotherapists). Only two contrasts reached significance: emotional empathy (β = 0.59, p = .02) and motivational empathy (β = 0.61, p = .02). Scientific quality (p = .10) and cognitive empathy (p = .08) did not. Identification sat at chance, at 45 % accuracy (χ2 test, p = .27).
The design is a cross-sectional, pre-registered comparison that the authors call explorative. Reader questions and expert answers from 26 "Dagens Nyheter" columns (2020 to 2024) were paired with answers from a GPT-4 agent that used retrieval augmentation over 20 of those columns and was instructed to mimic a professional psychologist. The research team edited the generated answers "to remove specific signs of AI writing (like bullet points)" before pairing them. Empathy was a three-item scale covering emotional, cognitive and motivational components (Cronbach's alpha 0.89, N = 208). Scientific quality was a single item. Both were rated on a 1 to 5 scale. The authors suggest the nonsignificant results may reflect low sensitivity: the intended sample was not reached, and a post-hoc analysis shows the study could detect only moderate effects (OR ≥ 1.6–1.7). The nonsignificant contrasts are therefore underpowered, not evidence of equivalence.
This extends Can language models match therapist empathy in real conversations?, which found a similar single-turn advantage over trainee therapists on a behavioral activation task. The two share a boundary: each response is judged on its own. The clinician ratings here measure how a reply reads, not what a reader did with it. That is the gap Can language models truly understand therapeutic ruptures? exposes in the rupture setting, where a label match can hold while the underlying reading differs. The contrast with Does warmth training make language models less reliable? lies in what is measured. This study's quality item is one rater judgment. It cannot show the factual or safety degradation that the warmth note reports, so the two findings concern different failure surfaces.
The excerpt does not establish that people seeking help would be better served, and it reports no outcomes. The six test columns come from one newspaper. One expert author accounts for 53 % of all expert answers and 67 % of the test set. The excerpt says the columns were chosen for topical variety but does not show they represent wider advice. Clinicians rated written answers (the expert answers had a median of 834 words), not conversations, and the excerpt is silent on multi-turn use. The edits to the AI text mean the rated text was not raw model output, and the excerpt does not say how much the edits changed it. The defensible reading is narrow. In this written format, with these edits and this sample, clinicians did not detect AI authorship and rated the AI text at least as well on empathy and soundness. The finding does not show that AI can stand in for a clinician.
Inquiring lines that read this note 21
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can artificial systems establish authority in domains requiring expert judgment? How do clinicians calibrate trust in AI medical recommendations?- How much do edited AI responses versus raw outputs affect clinician ratings?
- Are newspaper advice columns representative of broader professional psychological guidance?
- Do patients show the same bias toward expert-labeled medical advice?
- Why do clinicians fail to act on correct AI suggestions in real care?
- What role does interface design play in clinician adoption of AI tools?
- What evidence would prove medical AI actually works in clinics?
- Do patients actually perceive AI as worse at addressing their unique medical needs?
- Why did primary care physicians review only 73% of AI-generated transcripts?
- Can clinicians reliably distinguish high-quality AI advice from low-quality advice by appearance alone?
- Do physicians follow incorrect advice more when they trust its source?
- How much diagnostic accuracy is gained when physicians receive expert advice?
- Can algorithmic aversion explain clinicians' skepticism of AI recommendations?
- Do consensus criteria identify behaviors where physicians and models differ most?
- Does AI change clinician cognition or just increase reliance on predictions?
- Why does labeling advice as AI from a doctor change how people trust it?
- Can people tell which medical advice is accurate based only on how it reads?
- How does trusting wrong AI advice change what medical action people decide to take?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can language models match therapist empathy in real conversations?
Do LLMs' high empathy scores on isolated responses translate to therapeutic skill in actual ongoing treatment? This explores whether single-turn advantage predicts real-world therapeutic performance.
same single-turn pattern of advantage, here with clinicians rating newspaper advice, and the same isolated-response boundary
-
Can language models truly understand therapeutic ruptures?
When LLMs match expert labels for therapeutic ruptures, are they demonstrating genuine clinical understanding or relying on surface-level linguistic patterns? This matters because high identification scores may mask fundamentally different reasoning.
a rating of how a reply reads does not show what the reader did with it
-
Does warmth training make language models less reliable?
Explores whether training models for empathy and warmth creates a hidden trade-off that degrades accuracy on medical, factual, and safety-critical tasks—and whether standard safety tests catch it.
contrast: a single rater-judged quality item is not an accuracy test
-
Does the label on advice shape how clinicians judge it?
When clinicians believe advice comes from an expert, do they rate it higher regardless of who actually wrote it? This matters because it reveals whether judgments track the advice itself or just its claimed source.
the sibling finding: ratings moved with believed authorship, which this parity result did not test
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Artificial intelligence vs. human expert: Licensed mental health clinicians' blinded evaluation of AI-generated and expert psychological advice
- Do as AI say: susceptibility in deployment of clinical decision-aids
- People Overtrust AI-Generated Medical Advice despite Low Accuracy
- People Defer to AI Moral Advice, But Not Blindly
- Large Language Models Do Not Simulate Human Psychology
- Evaluating Large Language Models in Theory of Mind Tasks
- VCounselor: A Psychological Intervention Chat Agent Based on a Knowledge-Enhanced Large Language Model
- Comparing Human and AI Therapists in Behavioral Activation for Depression: Cross-Sectional Questionnaire Study
Original note title
clinicians rated GPT-4 advice as equal or more empathetic and scientifically sound than expert advice and could not reliably tell which was which