INQUIRING LINE

Doctors said AI advice looked less trustworthy when labeled 'AI' — but still followed it just as often, even when wrong.

Can annotation and explanation labels reduce automation bias in clinical settings?

This explores whether tagging AI output in clinical settings, either by labeling where advice came from or by attaching explanations, makes clinicians less likely to follow the AI when it is wrong.


This explores whether labels and explanations attached to AI advice help clinicians catch the AI's mistakes instead of going along with them. The short answer from this collection is that labels change what clinicians say about AI advice more than what they do with it. The most direct evidence comes from a study of radiologists Does labeling advice as AI change how clinicians use it?. They saw identical advice labeled either 'AI' or 'human expert.' They rated the AI-labeled version lower, but their diagnostic accuracy followed whether the advice was correct, not which label it carried. Clinicians who distrusted AI advice on paper still followed it in practice, including when it was wrong. That is what automation bias looks like, and a source label didn't fix it.

If labeling the source doesn't help, changing what the AI hands over might. The 'learning to guide' approach Can AI guidance reduce anchoring bias better than AI decisions? stops giving the clinician an answer to accept or reject. Instead the machine points out which parts of the input deserve attention and leaves the judgment to the human. Without a ready-made verdict to anchor on, anchoring bias goes away. The lesson is that automation bias may come less from how advice is labeled and more from the fact that a finished decision is being offered at all.

Explanations come with a catch of their own. Research on fine-tuning Does supervised fine-tuning improve reasoning or just answers? found that a model can become more accurate while its step-by-step reasoning gets worse, so it ends up justifying answers after the fact. An explanation attached to clinical advice could therefore sound convincing without showing how the answer was actually reached, and that could make over-trust worse rather than better. A related warning Can AI models be truly free from human bias? is that high headline accuracy can hide many confident errors. A '95% accurate' label still leaves a lot of wrong calls at scale.

One kind of label does show measurable gains, but it is applied to the model rather than to the clinician. When assistants were given an explicit list of what they don't know about the user, harmful advice and sycophancy fell by 50 to 75 percent and hallucinations by about half Do language models know what they don't know about users?. In a clinical setting, flagging what the AI hasn't seen might help more than flagging who produced the advice.

The collection doesn't have a head-to-head clinical trial comparing explanation formats against automation bias, so treat this as converging hints rather than a settled answer. Taken together, they suggest that labels and explanations alone don't do much, and that changing the role the AI plays, from deciding to guiding, may matter more.


Sources 5 notes

Does labeling advice as AI change how clinicians use it?

Radiologists rated AI-labeled advice lower than identical advice labeled human-expert, yet their diagnostic accuracy depended on whether the advice was correct, not its source. This suggests labels shape what clinicians think about advice but not how they use it.

Can AI guidance reduce anchoring bias better than AI decisions?

Learning to Guide eliminates anchoring bias and unassisted hard cases by having machines supply interpretive guidance rather than autonomous decisions, keeping responsibility with humans while improving their judgment through enhanced perception.

Does supervised fine-tuning improve reasoning or just answers?

Supervised fine-tuning improves final-answer accuracy on benchmarks but cuts Information Gain by 38.9 percent, meaning models generate correct answers through post-hoc rationalization rather than genuine inferential steps. Standard metrics miss this degradation because they only measure final correctness.

Can AI models be truly free from human bias?

Research shows that 'theory-free' AI models mask bigotry behind high accuracy metrics while committing fundamental statistical errors. A 95% accurate criminal justice system would wrongly convict thousands, demonstrating that model sophistication does not validate causal inference.

Do language models know what they don't know about users?

Research shows assistants suffer from sycophancy and hallucination because they have no representation of what remains unknown about users. Adding a schema of labeled unknowns to prompts reduced harmful advice and sycophancy by 50–75% and cut hallucination rates by roughly half.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.