SYNTHESIS NOTE
Topics›Domain Specialization›this note

How much does wrong AI advice harm radiologist accuracy?

When mammography radiologists receive incorrect AI suggestions labeled as system output, how much does their diagnostic accuracy decline? This matters for understanding automation bias in clinical workflows.

Synthesis note · 2026-10-06 · sourced from Domain Specialization

Dratsch et al. report, through RSNA, that radiologists reading mammograms with AI decision support became significantly worse at assigning the correct BI-RADS category when the suggestion was wrong. In a prospective experiment, 27 radiologists read 50 mammograms and gave their BI-RADS assessments with AI assistance, using two randomized sets: a training set of 10 with correct AI suggestions, and a test set of 40 in which 12 carried incorrect categories, purportedly suggested by AI. Inexperienced radiologists assigned the correct score "in almost 80% of cases" when the suggestion was correct and fell to less than 20% when it was wrong. Experienced radiologists, with more than 15 years of experience on average, dropped from 82% to 45.5%.

The excerpt reads the result as automation bias, defined as "the tendency of humans to favor suggestions from automated decision-making systems." It compares cases where the purported suggestion was right with cases where it was wrong, and the drop appears in both experience groups, though the experienced drop is smaller. The authors tie the concern to the workflow itself: "Given the repetitive and highly standardized nature of mammography screening, automation bias may become a concern when an AI system is integrated into the workflow." The excerpt notes that earlier studies of computer-aided detection had found performance impairments, but none had looked at AI systems and accurate readings. The safeguards it lists (showing confidence or per-output probabilities, teaching users how the system reasons, and keeping users accountable for their own decisions) are offered as possibilities, not tested.

Against the library, this is the measured side of a risk that Can AI guidance reduce anchoring bias better than AI decisions? describes in design terms. That note argues that deferral-style systems risk anchoring bias, where the human over-trusts the machine's decision. The mammography excerpt gives that concern an empirical case in a clinical reading task, though it names the mechanism automation bias rather than anchoring, and it does not test the guidance-style design the LTD-LTG note proposes as the fix. The closer link on the cognitive side is Does AI assistance actually harm the way developers learn?, where high-engagement interaction patterns preserved learning outcomes. That fits the excerpt's accountability safeguard, but the two studies measure different things (conceptual learning in a coding session versus accuracy on individual readings), so the connection is a hypothesis rather than a shared finding.

The excerpt is a press release, not the paper, so sample details beyond the headline counts, the statistical tests behind "significantly worse," and the per-reader spread are not given here. "Purportedly" carries the most weight: the wrong categories were labeled as AI output, so the study tests how readers respond to a labeled suggestion that is wrong, not how a deployed system's actual errors would affect them. Twenty-seven readers working through a fixed set of 50 cases is also a long way from a live screening workflow. The excerpt does not show why readers deferred either; the researchers plan eye-tracking to study that, so automation bias remains their interpretation. The implication is narrower than the headline. The result supports the authors' call for safeguards when AI enters the reading workflow, but the excerpt does not show that confidence displays, reasoning explanations or accountability prompts would reduce the effect.

Inquiring lines that read this note 15

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do clinicians calibrate trust in AI medical recommendations?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 95 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

Dratsch et al. find wrong BI-RADS categories labeled as AI output cut accuracy for experienced radiologists from 82% to 45.5%