MedGemma Technical Report
Artificial intelligence (AI) has significant potential in healthcare applications, but its training and deployment are challenging due to healthcare’s diverse data, complex spectrum of possible tasks, and the need to preserve privacy. Foundation models that perform well on various medical tasks and require less task-specific tuning data are critical to accelerating the development of AI for healthcare applications. In this technical report, we introduce MedGemma, a new collection of medical vision–language foundation models based on Gemma 3 4B and 27B. MedGemma demonstrates advanced medical understanding and reasoning on images and text, significantly exceeding the performance of similar-sized generative models and approaching the performance of task-specific models, while maintaining the general capabilities of the Gemma 3 base models. For out-of-distribution tasks, MedGemma achieves 2.6-10% improvements on medical multimodal question answering, 15.5-18.1% improvements on chest X-ray finding classification, and 10.8% improvement on agentic evaluations compared to the base models. Fine-tuning MedGemma further improves performance in subdomains, reducing errors in electronic health record information retrieval by 50% and reaching comparable performance to existing specialized state-of-the-art methods for pneumothorax classification and histopathology patch type classification. We additionally introduce MedSigLIP, a medically-tuned vision encoder derived from SigLIP. MedSigLIP powers the visual understanding capabilities of MedGemma and, as an encoder, it achieves performance comparable to or better than specialized medical image encoders. Taken together, the MedGemma collection provides a strong foundation of medical image and text capabilities, with potential to significantly accelerate medical research and development of downstream applications. More details about the MedGemma collection, including tutorials and instructions for downloading the model weights, can be found at https://goo.gle/medgemma.
Introduction. The landscape of modern healthcare is characterized by the generation and use of an unprecedented volume and diversity of data. Diagnosis, treatment, and monitoring rely on synthesizing information from disparate sources and specialties. Recently developed large multimodal models (LMMs), trained on massive and diverse datasets, exhibit remarkable capabilities in detecting complex patterns, generating coherent text, and processing visual information (Achiam et al., 2023; Alayrac et al., 2022; Chen et al., 2022; Liu et al., 2023, 2024; OpenAI, 2023; Touvron et al., 2023). These capabilities mark a potential paradigm shift in assisting with current workflows and extracting novel insights.
While general-purpose (non-medically tuned) LMMs demonstrate impressively broad abilities, generic models can lack nuanced medical understanding and the ability to interpret and reason about medical data in a robust way (Han et al., 2023; Labrak et al., 2024; Singhal et al., 2023b,c; Toma et al., 2023; Tu et al., 2024; Yang et al., 2024). Recognizing this gap, we created MedGemma, a new suite of open, medically-tuned, vision-language foundation models. These models represent the latest addition to the Health AI Developer Foundations (Kiraly et al., 2024) collection. Built upon the robust architecture of Gemma 3 (Gemma-Team et al., 2025), the MedGemma models are designed to interpret and reason about medical images and text while retaining the strong general-purpose capabilities present in Gemma 3.
In this report, we focus on two MedGemma models: a 4B variant that can accept text, images, or both as input, and a 27B variant that is optimized for text-only inputs. Both models output text. MedGemma 4B demonstrates strong performance on Vision Question Answering (VQA) benchmarks compared to prior SOTA models like Med-Gemini (Saab et al., 2024; Yang et al., 2024) despite being considerably smaller. Both MedGemma 4B and 27B are highly competitive on challenging A high level overview of the released models is shown in Fig. 1. More details about the MedGemma collection, including tutorials and links to download all of the above models, can be found at https: //goo.gle/medgemma.
Method. For general purpose data replay during pretraining, original data mixtures from SigLIP (Zhai et al., 2023) and Gemma 3 (Gemma-Team et al., 2025) were leveraged. The medical training and evaluation datasets largely followed the datasets in Med-Gemini (Yang et al., 2024). In this section, we outline the specific changes or differences in datasets relative to Med-Gemini.
Text-only datasets: For text datasets, we sampled responses and logits from a large IT (instructiontuned) teacher using the train splits of multiple medical QA datasets, including MedQA (Jin et al., 2021), MedMCQA (Pal et al., 2022), PubMedQA (Jin et al., 2019), MedExpQA (Alonso et al., 2024), AfriMed-QA (Olatunji et al., 2024), HealthSearchQA (Singhal et al., 2023a), and LiveQA (Abacha et al., 2017). We also sampled responses and logits for approximately 200,000 synthetic medical questions generated by asking the same large IT teacher to generate a new question using 5 randomly sampled questions from the above datasets as examples.
Multimodal datasets: Relative to Med-Gemini, the multimodal capabilities of MedGemma are currently focused on 2D medical images (e.g. X-ray, 2D slices from CT/MRI); 3D volumes and genomic datasets described in Yang et al. (2024) were not included. Additionally, we and others have identified potential data quality issues in PathVQA and MedVQA. Thus, we removed them from the training dataset. We did not include PAD-UFES-20 in the post-training dataset since it focuses on 6-class classification of very specific lesion types, which is not in line with the goal of more general purpose dermatology capabilities and use cases. For the PMC-OA component of the training data, we only included the single panel medical images from PMC-OA for better data quality. Relative to Med- Gemini we also introduced a larger internal collection for ophthalmology (184,852 more retinal fundus images), dermatology (51,049 more dermatology images with 210 different skin conditions), histopathology (a total of ∼32.5 million patch-text pairs), and radiology data (54,573 more CT 2D Our data preparation followed Yang et al. (2024) closely. Image padding and resizing algorithms remain the same, but because the vision encoder is different in Gemma 3, our images were resized to 896×896 instead of 768×768. Following Gemma 3, we use the SentencePiece tokenizer with 262,000 entries. Additionally, for CT images, we preselected three windows and converted them into the RGB color channels of the input image to highlight (1) bone and lung, window-width: 2250, window-level: -100; (2) soft tissue, window-width: 350, window-level: 40; (3) brain, window-width: 80, window-level: 40.
The MedGemma model architecture follows Gemma 3 (Gemma-Team et al., 2025) and is compatible with all existing Gemma infrastructure. The vision encoder for Gemma 3 is the 400M variant of the SigLIP encoder (Zhai et al., 2023) and is shared across the different Gemma language model sizes (4B, 27B). The input image resolution is 896×896 with pixel values normalized to [-1, 1]. The language model component also follows Gemma 3, featuring arbitrary image-text interleaving and long context (128k). Similar to Gemma 3, MedGemma was trained on TPUv4, TPUv5e, and TPUv5p, leveraged pre-computed visual tokens for memory saving, and used data and model shardings for multi-pod training.
The MedGemma 4B multimodal model utilized all of the following steps while the text-only version of MedGemma 27B leveraged the post-training stage alone.
Vision Encoder Enhancement for MedGemma: To improve the vision encoder’s capability of encoding and distinguishing subtle differences in medical images, we fine-tuned the vision encoder in Gemma 3 (SigLiP-400M) using over 33M medical image-text pairs (635k from various medical modalities and 32.6M histopathology patches) as listed in Table 1. To retain SigLIP’s existing performance, its original training data (e.g., WebLI) were retained and medical data was mixed with 2% weight into the training. While the Gemma 3 vision encoder works with 896×896 resolution, we found that many medical vision tasks worked reasonably well at 448×448 resolution (Table 15).
Discussion. We introduced MedGemma, a new collection of medical vision-language foundation models and MedSigLIP, a multi-domain medical image encoder. These models were built upon Gemma 3, with optimization for medical domains. We evaluated across a range of medical benchmarks across clinical reasoning, biomedical knowledge, report generation, and medical image classification, finding strong performance for MedGemma and MedSigLIP. Performance improved further after fine-tuning, highlighting the potential for these open models to be used as a starting point for developing useful AI applications for healthcare.
With an increasing number of options available to developers building AI applications in healthcare, MedGemma provides specific advantages over general models. These advantages are largely due to optimized incorporation of domain specific data for both pre-training and post-training and are illustrated by the improvements over base Gemma 3 models across all benchmarks evaluated and the achievement of performance on par with much larger models.
When compared to general API-based models like Gemini, MedGemma is likely the preferred model if the use case requires any of the following: a frozen model for documentation and reliability, sensitivity to training or inference costs, ability to run locally or offline, specific medical image and text capabilities, or full control over model adaptation. Large models like Gemini remain a viable choice where the user requires optimal broad performance without the above constraints, and large models may additionally be used in concert with models like MedGemma in agentic settings.
The MedGemma collection of models enables a wide range of potential downstream applications for the developer community. The multimodal capabilities, including access to image and text embeddings, may be particularly useful for medical image retrieval. This could aid in interpretation by referencing similar past cases as well as enabling development of research cohorts and creating educational tools. MedGemma allows for the integration of diverse data, linking radiology, histopathology, ophthalmology, and dermatology images with clinical information. The specialized text capabilities of the models can also extract key concepts from imaging reports and clinical notes, streamlining tasks such as matching patients for clinical trials, conducting pharmacovigilance reviews, or analyzing healthcare quality metrics. The models’ ability to understand medical images and generate reports can also be fine-tuned to better assist radiologists and other clinicians in their workflow and improve how findings are communicated to patients. In addition to standalone use, these models can also serve as powerful tools within agentic frameworks, combining abilities across different modalities for customized and comprehensive solutions.
In this report, we evaluated the performance of MedGemma and MedSigLIP on a broad set of established benchmarks in order to provide a snapshot of the model capabilities. However, we note that limitations exist for these benchmarks. For one, automated benchmarks represent only the first step towards validating real-world utility (Alaa et al., 2025; Mahmood, 2025). Additionally, some benchmarks may be near saturation in terms of model performance, with minimal headroom for improvements, thus hindering the measurement of progress. As such, further work is warranted to continue evaluation of these models on new, high quality (and more challenging) benchmarks aimed at better reflecting real-world utility (e.g. Bedi et al., 2025). More work is also needed to understand the performance capabilities and requirements in regard to actual application development, including their incorporation into agentic frameworks.
Conclusion. In this work, we showed that MedGemma models demonstrate robust capabilities across a variety of vision-language and text-only medical tasks. We also showed that MedSigLIP demonstrates robust multi-domain capabilities, and can thus serve as a strong medical foundation model. The breadth and efficiency of these models offers exciting possibilities to address a range of use cases. At the same time, thoughtful validation of safety, performance, and reliability for any downstream applications remains a critical aspect to advance the use of multimodal AI models in medicine. By providing these MedGemma and MedSigLIP models to the developer community with a permissive license, we hope to see them enable useful and innovative medical applications.
Lines of inquiry this paper opens 14
Research framings built by reading the notes related to this paper — the questions it feeds into.
How do curriculum design and feedback approaches affect model learning? Does scaling reasoning capability create fundamental tradeoffs in control and reliability?- Why does general reasoning not transfer to knowledge-intensive medical domains?
- Why do medical and mathematical tasks require fundamentally different model capabilities?
- Why do medical diagnoses require human judgment even with AI assistance?
- What makes reasoning auditable in medical AI decision support?
- Why does medical knowledge require continuous access to current sources?
- Does medical AI accuracy depend more on knowledge or reasoning ability?
- Does medical domain competency require knowledge injection or better prompting?
- Can medical diagnosis depend less on knowledge and more on orchestration?
- Does this colonoscopy finding apply to other medical specialties using AI?
- Can people tell which medical advice is accurate based only on how it reads?
- How does blinded rating of diagnoses compare to real clinical outcomes?