Does the label on advice shape how clinicians judge it?
When clinicians believe advice comes from an expert, do they rate it higher regardless of who actually wrote it? This matters because it reveals whether judgments track the advice itself or just its claimed source.
Clinicians' judgments followed the label on the advice more than the advice's actual author. When a reply was perceived as expert-written, it scored higher on scientific quality and on all three empathy components, and it was preferred. The excerpt reports coefficients of −1.89 for scientific quality, −3.62 for emotional empathy, −3.02 for cognitive empathy and −2.74 for motivational empathy, all at p < .001. On preference, actual authorship (β = 6.96, p = .002) and perceived authorship (β = 6.26, p = .001) were both significant. Participants showed a "93.55 % preference for perceived expert advice," and they chose it "regardless of whether the answer was actually authored by AI or an expert."
The excerpt's own term is influence: "perceived authorship influenced ratings." Its highlights name "potential biases in the acceptance of AI-generated mental health support." The mechanism it offers is belief about authorship, not the text itself. The identification result sharpens this. Clinicians were at chance (45 % accuracy), so the beliefs that moved their scores were not reliable readings of who had written the text. The excerpt also reports a significant interaction between perceived and actual authorship (β = −12.29, p = .001), shown in its Fig. 1. The text does not unpack that interaction beyond the 93.55 % figure. The sign convention for the rating coefficients is not explained, so this note reports direction only. Perceived authorship was a rater's reported guess, not an assigned condition, so the analysis describes an association between belief and score.
This is the bias the paper names, and it qualifies the parity result in Can clinicians tell GPT-4 advice apart from expert advice?. That parity was measured with authorship hidden from the rater's judgment. The excerpt does not test what happens when the label is visible, so parity cannot be assumed to survive disclosure. The pattern also echoes Can language models truly understand therapeutic ruptures?, where a match to a reference label can hide a different method underneath. Here the rater's score tracks a label (who wrote the text) rather than the text. Both cautions point at the same evaluation risk: a score can faithfully reflect a label while telling you little about the thing the label names.
The excerpt does not show whether disclosure would change the result, because it does not state the order in which the authorship guess and the ratings were collected. It does not show that patients or other lay readers would show the same pattern. The raters were 43 licensed clinicians, 40 of them psychologists. The practical question is open. Before disclosure rules are set for AI-written mental health advice, ratings would need to be compared with authorship shown and hidden. The excerpt does not report that comparison. The bias is best read as a documented tendency in one blinded study, not as a measured effect size for practice.
Inquiring lines that read this note 9
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can artificial systems establish authority in domains requiring expert judgment? How do clinicians calibrate trust in AI medical recommendations?- Do patients show the same bias toward expert-labeled medical advice?
- Would clinicians' ratings change if authorship was visible from the start?
- Why did clinicians guess authorship at chance level despite strong preferences?
- Do physicians follow incorrect advice more when they trust its source?
- How much diagnostic accuracy is gained when physicians receive expert advice?
- Why does labeling advice as AI from a doctor change how people trust it?
- Can people tell which medical advice is accurate based only on how it reads?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can clinicians tell GPT-4 advice apart from expert advice?
This study explores whether trained clinicians can distinguish AI-generated psychological advice from expert advice, and how they rate the quality and empathy of each. The question matters for understanding whether AI might reliably supplement human expertise in mental health settings.
the sibling parity finding; this note shows that believed authorship moved the scores, which the blinded parity result did not test
-
Can language models truly understand therapeutic ruptures?
When LLMs match expert labels for therapeutic ruptures, are they demonstrating genuine clinical understanding or relying on surface-level linguistic patterns? This matters because high identification scores may mask fundamentally different reasoning.
a label match can hide a different method, and here a rater's score tracks the authorship label rather than the text
-
Do reader judgments reflect actual authorship or just their beliefs?
When readers evaluate research abstracts, do their ratings track who actually wrote them, or are they shaped by what they believe about authorship—even when those beliefs are wrong?
Evidence for: reader ratings of abstracts track authorship beliefs even when those beliefs are wrong, as A's belief-driven preference does
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Artificial intelligence vs. human expert: Licensed mental health clinicians' blinded evaluation of AI-generated and expert psychological advice
- Do as AI say: susceptibility in deployment of clinical decision-aids
- People Overtrust AI-Generated Medical Advice despite Low Accuracy
- People Defer to AI Moral Advice, But Not Blindly
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
- On Information Distortions in Online Ratings
- Consult Evaluation: Scottish Government's Non-surgical Cosmetic Procedures Consultation
- The human-authorship halo: attribution bias in literary style evaluation by humans and AI
Original note title
perceived expert authorship shaped clinicians' preferences — 93.55 % went to advice believed expert-written, whatever its actual author