SYNTHESIS NOTE
Topics›Domain Specialization›this note

Can clinicians tell GPT-4 advice apart from expert advice?

This study explores whether trained clinicians can distinguish AI-generated psychological advice from expert advice, and how they rate the quality and empathy of each. The question matters for understanding whether AI might reliably supplement human expertise in mental health settings.

Synthesis note · 2026-10-06 · sourced from Domain Specialization

In a blinded comparison, licensed clinicians rated GPT-4 advice as equal or more favorable than expert advice from a Swedish newspaper advice column, on empathy and scientific quality, and they could not reliably tell the two apart. The excerpt reports that "AI-generated advice received equal or more favorable ratings across all measures" over 104 response pairs rated by 43 clinicians (40 psychologists, three psychotherapists). Only two contrasts reached significance: emotional empathy (β = 0.59, p = .02) and motivational empathy (β = 0.61, p = .02). Scientific quality (p = .10) and cognitive empathy (p = .08) did not. Identification sat at chance, at 45 % accuracy (χ2 test, p = .27).

The design is a cross-sectional, pre-registered comparison that the authors call explorative. Reader questions and expert answers from 26 "Dagens Nyheter" columns (2020 to 2024) were paired with answers from a GPT-4 agent that used retrieval augmentation over 20 of those columns and was instructed to mimic a professional psychologist. The research team edited the generated answers "to remove specific signs of AI writing (like bullet points)" before pairing them. Empathy was a three-item scale covering emotional, cognitive and motivational components (Cronbach's alpha 0.89, N = 208). Scientific quality was a single item. Both were rated on a 1 to 5 scale. The authors suggest the nonsignificant results may reflect low sensitivity: the intended sample was not reached, and a post-hoc analysis shows the study could detect only moderate effects (OR ≥ 1.6–1.7). The nonsignificant contrasts are therefore underpowered, not evidence of equivalence.

This extends Can language models match therapist empathy in real conversations?, which found a similar single-turn advantage over trainee therapists on a behavioral activation task. The two share a boundary: each response is judged on its own. The clinician ratings here measure how a reply reads, not what a reader did with it. That is the gap Can language models truly understand therapeutic ruptures? exposes in the rupture setting, where a label match can hold while the underlying reading differs. The contrast with Does warmth training make language models less reliable? lies in what is measured. This study's quality item is one rater judgment. It cannot show the factual or safety degradation that the warmth note reports, so the two findings concern different failure surfaces.

The excerpt does not establish that people seeking help would be better served, and it reports no outcomes. The six test columns come from one newspaper. One expert author accounts for 53 % of all expert answers and 67 % of the test set. The excerpt says the columns were chosen for topical variety but does not show they represent wider advice. Clinicians rated written answers (the expert answers had a median of 834 words), not conversations, and the excerpt is silent on multi-turn use. The edits to the AI text mean the rated text was not raw model output, and the excerpt does not say how much the edits changed it. The defensible reading is narrow. In this written format, with these edits and this sample, clinicians did not detect AI authorship and rated the AI text at least as well on empathy and soundness. The finding does not show that AI can stand in for a clinician.

Inquiring lines that read this note 21

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can artificial systems establish authority in domains requiring expert judgment? How do clinicians calibrate trust in AI medical recommendations? Can real-time working alliance measurement improve therapy outcomes? Can AI systems perform peer review as effectively as humans? How do users confuse explanation quality with actual system accuracy?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 93 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

clinicians rated GPT-4 advice as equal or more empathetic and scientifically sound than expert advice and could not reliably tell which was which