Does medical pretraining actually improve model performance?
When medical models are fairly compared to their base counterparts with per-model prompt tuning and statistical significance testing, do they consistently outperform on medical question answering tasks?
The paper argues that medical domain-adaptive pretraining (DAPT), the common practice of continuing to pretrain an open general-domain model on PubMed and medical textbooks, does not reliably make models better at medical question answering. Comparing seven medical LLMs and two medical VLMs against their own base models, it finds that "all medical VLMs and nearly all medical LLMs fail to consistently improve over their general-domain counterparts" in zero- and few-shot prompting. The headline figure is for the 3-shot setting: across tasks and model pairs, medical LLMs beat their base models in 12.1% of cases, tie in 49.8%, and are significantly worse in the remaining 38.2%. The paper's closing claim is the stronger reading: "state-of-the-art general-domain models may already exhibit strong medical knowledge and reasoning capabilities."
The mechanism is comparison design, not a new benchmark. The paper makes three moves: it compares each medical model head-to-head with the base model it was derived from, so that only the medical pretraining differs; it selects prompt format and few-shot examples separately for each model on the validation set; and it reports bootstrapped confidence intervals on the test set. The second move matters because the paper notes that the "optimal" prompt choice "rarely correlates between different models," so a single prompt tuned for the medical model can flatter it. The zero-shot numbers show the size of the effect. When the prompt is optimized only for the medical model and uncertainty is ignored, medical LLMs and VLMs outperform their base models on average in 70.5% and 62.5% of QA tasks. After per-model prompt selection and significance testing, the improvements are significant in only 9.4% and 6.3% of tasks. The paper says its ablations show these practices "substantially impact conclusions."
Against the nearby notes, this is a measured test of one route in the knowledge-injection taxonomy. How do knowledge injection methods trade off flexibility and cost? places domain expertise embedded in weights during pretraining or fine-tuning on the static side, where the cost is training and the payoff is inference speed. Continued pretraining on medical text is that route, and this paper finds its payoff over the base model is largely absent in closed-ended QA once comparisons are fair. It also qualifies the knowledge-dominance reading in Does medical AI need knowledge or reasoning more?. Medicine may be knowledge-dominant, but if general models already hold that knowledge, the gap may lie in eliciting it rather than injecting it. The two readings are compatible: the paper's own phrasing is that general models can be "leveraged effectively when prompted appropriately."
The excerpt does not establish more than this. It covers closed-ended QA scored by exact match under greedy decoding. The authors concede that open-ended consumer health QA and information extraction from clinical notes are untested, and that "it is certainly possible that an analysis like ours would find improved performance on such tasks." The model set is not exhaustive, and the excerpt does not show the per-pair results behind the BioMistral-7B exception. The implication, at the strength the evidence allows: on closed-ended medical QA, a medical checkpoint should not be credited over its own base model until both are compared under per-model prompts with uncertainty reported. The result does not show that DAPT is useless for clinical extraction or for open-ended tasks.
Inquiring lines that read this note 5
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do curriculum design and feedback approaches affect model learning? How do clinicians calibrate trust in AI medical recommendations? How does fine-tuning trade off accuracy against reasoning quality?Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
How do knowledge injection methods trade off flexibility and cost?
When and how should domain knowledge enter an AI system? This explores the speed, training cost, and adaptability trade-offs across four injection paradigms, and when each approach suits different deployment constraints.
this paper tests the static-embedding route the taxonomy describes, finding its gain largely disappears in a matched comparison
-
Does medical AI need knowledge or reasoning more?
Medical and mathematical domains may require fundamentally different AI training priorities. If medical accuracy depends primarily on factual knowledge while math depends on reasoning quality, should we build and evaluate these systems differently?
the claim that general models already hold medical knowledge suggests the gap may be elicitation, not injection
-
Does medical model architecture or training data drive performance gains?
MedGemma claims its medical improvements come from domain-specific training data rather than architectural changes. But the evidence comes only from the developers' own benchmarks, without independent validation or real-world clinical testing.
Contradicts A's null-gain finding: MedGemma's own evaluations show 2.6-10% and 15.5-18.1% gains from domain data over base Gemma 3, same architecture
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Medical Adaptation of Large Language and Vision-Language Models: Are We Making Progress?
- Capabilities of Gemini Models in Medicine
- HealthBench: Evaluating Large Language Models Towards Improved Human Health
- Clinical knowledge in LLMs does not translate to human interactions
- Knowledge or Reasoning? A Close Look at How LLMs Think Across Domains
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts
- People Overtrust AI-Generated Medical Advice despite Low Accuracy
- Thinking LLMs: General Instruction Following with Thought Generation
Original note title
medical domain-adaptive pretraining largely disappears once each model gets its own prompt and statistical uncertainty is counted