Doctors who got used to AI found fewer polyps once it was switched off. Does the same happen in other specialties?
Does this colonoscopy finding apply to other medical specialties using AI?
This explores whether the colonoscopy result (doctors who got used to AI polyp detection became worse at finding adenomas once the AI was switched off) also shows up in radiology, primary care, and diagnostic reasoning, or whether it is specific to endoscopy.
This explores whether the colonoscopy 'deskilling' finding carries over to other specialties that use AI. The short answer is that the corpus has no direct replication. No other study here measures what happens to clinicians' unassisted skill after months of working with AI. The Polish study is still unusual: across four centres, adenoma detection in standard colonoscopies fell from 28.4% to 22.4% once AI assistance had become routine Does AI polyp detection weaken endoscopists' unassisted performance?. It measured doctors after the AI was gone. Nearly every other medical AI study measures doctors while the AI is present. That gap is the first thing to know: the evidence we'd need to answer the question mostly hasn't been collected.
What other specialties do show is a related but different problem, which is how AI bends judgment in the moment. In mammography, wrong BI-RADS suggestions (the standard breast-imaging risk categories) cut experienced radiologists' accuracy from 82% to 45.5%, and inexperienced readers fell below 20% How much does wrong AI advice harm radiologist accuracy?. That is automation bias, which is not the same as skill decay, but the two can feed each other. If you defer to the machine often enough, you may stop practising the looking that the machine now does for you. The pattern runs in the opposite direction too. Another radiology experiment found that AI predictions didn't improve radiologists on average, because they underweighted the AI and wrongly treated their own read as independent of it Why don't radiologists benefit from AI predictions?. A related study found that radiologists rated advice lower when it was labelled 'AI', yet their accuracy tracked whether the advice was correct, not where it came from Does labeling advice as AI change how clinicians use it?. Together these suggest clinicians are steered by AI more than they realise or say, and that is the condition under which quiet skill erosion would be most likely.
One way to judge transfer is to ask what kind of skill is at risk. Colonoscopy is a perceptual skill: scanning the bowel and noticing a flat, pale lesion. That resembles what one note calls the expert's ability to choose 'which differences make a difference', a selective kind of observation that AI pattern-matching imitates without performing Can AI distinguish which differences actually matter?. Mammography, dermatology and pathology rely on the same kind of trained vision, so they are the most plausible places for the colonoscopy effect to repeat. Diagnostic reasoning is a different case. Medical AI performance tracks knowledge correctness more than reasoning quality Does medical AI need knowledge or reasoning more?. So what an internist could lose is probably breadth of recall rather than perceptual vigilance.
That makes the run of diagnostic-LLM successes look different. LLM help raised clinicians' top-10 differential accuracy from 36.1% to 51.7% on hard NEJM cases, mostly by widening the list of possibilities they considered Does LLM assistance help clinicians build better differentials?. AMIE beat primary care physicians on 28 of 32 specialist-rated axes, with its edge in inference rather than in taking the history Can an AI system diagnose better than primary care doctors?. AMIE also took 100 real urgent-care histories without a single safety stop Can conversational AI safely take patient histories without supervision?. Separately, another LLM outperformed hundreds of physicians on diagnosis and triage Can language models reason better than physicians at diagnosis?. Each of these is a deskilling risk in disguise. If the tool reliably supplies the breadth of the differential, physicians have less reason to build it themselves. None of these studies tests what happens when the tool is unavailable.
So the honest reading is this: the colonoscopy finding probably generalises most to image-based specialties, plausibly to diagnostic reasoning, and is untested almost everywhere. The question worth carrying away is not 'is AI accurate?' but 'what happens on the day it's down?' Medical AI evaluation is currently built to answer the first question and almost never the second.
Sources 10 notes
A Polish study of four endoscopy centers found adenoma detection rates fell from 28.4% to 22.4% in standard colonoscopies performed after clinicians began using AI assistance, suggesting continuous AI exposure may impair unaided diagnostic performance.
A 27-radiologist study found that incorrect BI-RADS suggestions caused experienced radiologists to drop from 82% to 45.5% accuracy, while inexperienced readers fell from nearly 80% to below 20%, demonstrating automation bias in mammography screening.
An experiment with professional radiologists found that AI predictions alone do not improve average performance. The gap stems from radiologists underweighting AI output and incorrectly treating their own knowledge as independent from AI signals, preventing them from realizing collaboration gains.
Radiologists rated AI-labeled advice lower than identical advice labeled human-expert, yet their diagnostic accuracy depended on whether the advice was correct, not its source. This suggests labels shape what clinicians think about advice but not how they use it.
Experts observe by choosing which differences matter (qualitative judgment); AI finds patterns and probabilities (quantitative). AI generates text from prompts without observing context, audience needs, or knowledge states—producing fabrication that mimics observation's form without its epistemic process.
Show all 10 sources
The KI/InfoGain framework reveals that medical domain accuracy correlates more strongly with knowledge correctness than reasoning quality, while mathematical domains show the inverse pattern. This distinction has direct implications for which training strategies to prioritize in each domain.
In a study of 20 clinicians on 302 NEJM cases, those with LLM access achieved 51.7% top-10 accuracy versus 36.1% without it. The authors attribute the gain to the LLM's wider differential scope, making lists more comprehensive.
An LLM-based diagnostic system called AMIE exceeded primary care physician performance in text-based simulated consultations across 149 case scenarios, scoring higher on 28 of 32 specialist-rated dimensions. The advantage lay in inference from gathered information rather than in eliciting history.
A single-arm study found that AMIE, a conversational AI system, conducted real clinical histories from 100 patients without requiring a single safety intervention by human supervisors. Patient attitudes toward AI improved after the interaction, though management plans trailed physicians on practicality and cost.
In physician-adjudicated vignette experiments and an emergency room study, a large language model outperformed hundreds of physicians on differential diagnosis, reasoning, triage, and clinical management tasks across multiple touchpoints.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Towards Conversational Diagnostic AI
- Clinical knowledge in LLMs does not translate to human interactions
- People Overtrust AI-Generated Medical Advice despite Low Accuracy
- Automation Bias in Mammography: The Impact of AI BI-RADS Suggestions on Reader Performance
- Sequential Diagnosis with Language Models
- Capabilities of Gemini Models in Medicine
- Combining Human Expertise with Artificial Intelligence: Experimental Evidence from Radiology
- Do as AI say: susceptibility in deployment of clinical decision-aids