Superhuman performance of a large language model on the reasoning tasks of a physician
A seminal paper published by Ledley and Lusted in 1959 introduced complex clinical diagnostic reasoning cases as the gold standard for the evaluation of expert medical computing systems, a standard that has held ever since. Here, we report the results of a physician evaluation of a large language model (LLM) on challenging clinical cases against a baseline of hundreds of physicians. We conduct five experiments to measure clinical reasoning across differential diagnosis generation, display of diagnostic reasoning, triage differential diagnosis, probabilistic reasoning, and management reasoning, all adjudicated by physician experts with validated psychometrics. We then report a real-world study comparing human expert and AI second opinions in randomly-selected patients in the emergency room of a major tertiary academic medical center in Boston, MA. We compared LLMs and board-certified physicians at three predefined diagnostic touchpoints: triage in the emergency room, initial evaluation by a physician, and admission to the hospital or intensive care unit. In all experiments—both vignettes and emergency room second opinions—the LLM displayed superhuman diagnostic and reasoning abilities, as well as continued improvement from prior generations of AI clinical decision support. Our study suggests that LLMs have achieved superhuman performance on general medical diagnostic and management reasoning, fulfilling the vision put forth by Ledley and Lusted, and motivating the urgent need for prospective trials. INTRODUCTION Artificial intelligence (AI) diagnostic support tools have been studied since the 1950s, following a landmark paper published in Science by Ledley and Lusted (1) who advocated for case-based benchmarks as an evaluation standard, a standard that has held for over the past half century (1–6). In particular, the New England Journal of Medicine clinicopathological case conference series has been seen an aspirational goal post, tested by every differential diagnosis generator from primitive Bayesian systems, symbolic rules-based systems, and natural-language symptom checkers.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How do clinicians calibrate trust in AI medical recommendations?- Can an AI system trained on text consultations handle diagnostic uncertainty in real patient encounters?
- How much does self-play training with LLM-simulated patients actually improve diagnostic accuracy?
- Does optimizing for differential diagnosis accuracy risk pushing AI systems toward premature problem-solving?
- Why do clinicians fail to act on correct AI suggestions in real care?
- How does expert annotation instability affect medical AI benchmarking?
- What evidence would prove medical AI actually works in clinics?
- Do patients actually perceive AI as worse at addressing their unique medical needs?
- How does blinded rating of diagnoses compare to real clinical outcomes?
- Does medical AI accuracy depend more on knowledge or reasoning ability?
- What prospective trials are needed to validate AI diagnostic claims?
- Does medical domain competency require knowledge injection or better prompting?
- How well do curated benchmark cases represent real clinical deployment?
- Can medical diagnosis depend less on knowledge and more on orchestration?
- How much diagnostic accuracy is gained when physicians receive expert advice?
- Does this colonoscopy finding apply to other medical specialties using AI?
- How much does prompt selection bias favor medical models over base models?
- Why did lay users with AI models fail to match unaided physicians on diagnosis?
- Do consensus criteria identify behaviors where physicians and models differ most?
- Can offline LLM evaluation predict performance in live clinical workflows?
- How much does missing images and tables limit LLM diagnostic reasoning?
- Can LLM performance on zebra cases predict results in routine clinical practice?
- Does medical fine-tuning help LLMs through knowledge or reasoning ability?
- How do LLM performances compare across different types of medical tasks?