Combining Human Expertise with Artificial Intelligence: Experimental Evidence from Radiology
Source: Agarwal, Moehring, Rajpurkar, Salz, NBER w31422 · 2026-08
Full automation using Artificial Intelligence (AI) predictions may not be optimal if humans have information not available to the AI (contextual information). We study human-AI collaboration using an information experiment with professional radiologists. Results show that providing (i) AI predictions does not improve performance on average, whereas (ii) contextual information does. Radiologists do not realize the gains from AI assistance because of errors in belief updating – they underweight AI predictions and treat their own information and AI predictions as statistically independent. Unless these mistakes can be corrected, the optimal human-AI collaboration design delegates cases either to humans or to AI, but rarely to AI assisted humans.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How do clinicians calibrate trust in AI medical recommendations?- Do radiologists' beliefs about AI-assisted performance match their actual outcomes?
- Does showing AI confidence scores reduce radiologist over-reliance on wrong suggestions?
- What safeguards help radiologists maintain independent judgment when using AI assistance?
- Why do clinicians fail to act on correct AI suggestions in real care?
- Why do experts resist AI recommendations that contradict their own judgments?
- How much diagnostic accuracy is gained when physicians receive expert advice?
- Does this colonoscopy finding apply to other medical specialties using AI?
- Why did lay users with AI models fail to match unaided physicians on diagnosis?
- Why do nurses misclassify emergencies differently with misleading AI assistance?
- Does AI change clinician cognition or just increase reliance on predictions?
- Why do radiologists fail to benefit from AI decision support?
- How does trusting wrong AI advice change what medical action people decide to take?
- Does optimizing for differential diagnosis accuracy risk pushing AI systems toward premature problem-solving?
- How does expert annotation instability affect medical AI benchmarking?
- Do patients actually perceive AI as worse at addressing their unique medical needs?
- What role does cost estimation play in steering diagnostic test ordering?
- Do physicians follow incorrect advice more when they trust its source?
- Why does labeling advice as AI from a doctor change how people trust it?
- Can an AI system trained on text consultations handle diagnostic uncertainty in real patient encounters?
- What evidence would prove medical AI actually works in clinics?
- Why did primary care physicians review only 73% of AI-generated transcripts?
- Can clinicians reliably distinguish high-quality AI advice from low-quality advice by appearance alone?
- Can annotation and explanation labels reduce automation bias in clinical settings?