Making an AI better at math doesn't make it better at medicine: math mostly breaks in the reasoning, medicine in the facts.
Why do medical and math domains need different types of model improvements?
This explores why the training approach that makes a model better at math (more and better step-by-step reasoning) doesn't carry over to medicine, and what kind of improvement medicine needs instead.
This explores why the training that makes a model better at math doesn't automatically make it better at medicine. The short answer from the corpus is that the two domains fail for different reasons. In math, models mostly go wrong in the reasoning: they set up the steps badly or lose the thread partway through. In medicine, they mostly go wrong on the facts: they recall a drug interaction incorrectly or misremember which symptom points to which condition. When researchers scored reasoning quality and knowledge correctness separately, medical accuracy tracked knowledge, while math accuracy tracked reasoning (Does medical AI need knowledge or reasoning more?). This explains a result that surprised many people: models trained to reason by imitating R1's reasoning traces didn't beat their own base models on medical tasks. Better reasoning was never the bottleneck (Why doesn't mathematical reasoning transfer to medicine?).
There may also be a physical explanation. One line of work suggests that facts are retrieved mostly in a network's lower layers, while reasoning adjustments happen in its higher layers (Why does reasoning training help math but hurt medical tasks?). If that's right, training aimed at reasoning works on the part of the model math depends on and mostly leaves the factual layers alone. Sometimes it can even disturb them. That fits a broader finding that extra reasoning isn't free: longer chains of thought can make a model talk itself out of a correct answer it had already reached (Why does more reasoning sometimes make models worse?). In math, a longer chain of reasoning is usually worth that risk. In medicine, each extra step is another chance to bring in a wrong fact.
The less obvious finding concerns what reinforcement learning actually does for medicine. You might expect it to teach the model new medical knowledge. According to Does RL improve domain reasoning by adding knowledge or removing it?, it mostly doesn't. Instead it rewards reasoning paths that avoid wrong facts, so the model learns to stop using what it misremembers. In medicine, then, RL helps by removing bad knowledge rather than adding new ability. That matters because models in specialized domains tend to be confidently wrong, and prompting tricks that help elsewhere don't fix this (Why do language models fail confidently in specialized domains?).
So is the answer simply "give medical models more medical data"? Partly, but that claim needs care. MedGemma attributes its gains over base Gemma 3 to curated medical data, with the architecture unchanged (Does medical model architecture or training data drive performance gains?). But a stricter evaluation found that once each model got its own optimized prompts and results came with confidence intervals, medical pretraining beat the base model on only about 9% of tasks (Does medical pretraining actually improve model performance?). Domain data matters, but much of the reported gain may come from how the comparison was set up.
The biggest medical gains may come from outside the model entirely. Wrapping o3 in a structured diagnostic process that orders tests, considers more diagnoses, and tracks cost raised accuracy and cut the cost per case by about 70%, and the improvement held when the same setup was used with other model families (Can orchestration strategies boost diagnostic AI without better models?). At the human end, models that correctly identified conditions about 95% of the time on their own helped members of the public get only about 34% right, no better than people without them (Why do LLMs fail when users interact with them?). In math, improving the model mostly means better reasoning. In medicine, improvement happens in three places: getting the model's facts right, the structured process around the model, and how people use what the model tells them.
Sources 10 notes
The KI/InfoGain framework reveals that medical domain accuracy correlates more strongly with knowledge correctness than reasoning quality, while mathematical domains show the inverse pattern. This distinction has direct implications for which training strategies to prioritize in each domain.
R1-distilled reasoning models fail to outperform base models on medical tasks because knowledge accuracy matters more than reasoning quality in medicine—the opposite of math. Fine-tuning cannot close this gap without domain-specific training data.
Two-phase inference model shows knowledge retrieval operates in lower network layers while reasoning adjustment happens in higher layers. This separation explains why reasoning training improves math but can degrade knowledge-intensive domains like medicine.
Tracking flip events shows that extra reasoning tokens don't just hit diminishing returns—they actively cause models to second-guess and overwrite previously-correct answers, making accuracy non-monotonic in trace length.
RL enhances medical reasoning by suppressing incorrect domain knowledge during reasoning—not by expanding what models know. Evidence shows RL achieves +12.4 point knowledge improvement by removing low-reward reasoning trajectories that invoke wrong facts.
Show all 10 sources
LLMs trained on general text lack sufficient exposure to domain-specific examples, leading to low accuracy paired with high confidence in clinical NLI tasks. Prompting techniques that improved general performance fail to reduce overconfidence in specialized domains.
MedGemma achieves 2.6-18.1% improvements over base Gemma 3 on medical tasks using identical architecture but curated medical data in pretraining and post-training. Gains include 50% error reduction in health record retrieval and performance comparable to specialized medical models.
Seven medical LLMs and two medical VLMs showed significant improvements over base models in only 9.4% and 6.3% of tasks respectively when prompts were optimized per model and confidence intervals were reported, down from 70.5% and 62.5% without these controls.
On 304 NEJM cases, MAI-DxO orchestration achieved 79.9% accuracy at $2,397 per case versus 78.6% at $7,850 for o3 alone. The gains transferred across model families, suggesting the benefit comes from the scaffold, not model weights.
A trial of 1,298 UK participants found GPT-4o, Llama 3, and Command R+ scored 94.9% accuracy alone but users achieved only 34.5% when identifying conditions, no better than controls. The gap lies in how users interpret and act on model suggestions, not model knowledge.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Knowledge or Reasoning? A Close Look at How LLMs Think Across Domains
- Medical Adaptation of Large Language and Vision-Language Models: Are We Making Progress?
- Capabilities of Gemini Models in Medicine
- Medical Reasoning in the Era of LLMs: A Systematic Review of Enhancement Techniques and Applications
- Sequential Diagnosis with Language Models
- HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
- Towards Accurate Differential Diagnosis with Large Language Models
- HealthBench: Evaluating Large Language Models Towards Improved Human Health