Foundation Models in Medicine โ 2025¶
๐๏ธ 2025 in Review
๐ Timeline¶
- 2025-01: Native-3D vision-only CT encoders arrive โ CT-FM (148k scans, intra-sample contrastive) and MEDFORM (CT+clinical-tabular contrastive) establish vision-centric, non-text 3D pretraining.
- 2025-02: Privacy-driven composition emerges โ MedForge (open-source-style LoRA merging), Med-LEGO (training-free model fusion), MedTok (graph-aware EHR tokenizer), and RELICT (memorization/replica auditing).
- 2025-03: Generalist-via-adaptation peaks โ continual-learning segmentation (CL-Net, 235 anatomies), unified ReID (MaMI), few-shot domain adaptation (MFM-DA), and the CUPCase finding that GPT-4o beats clinical LLMs on rare cases.
- 2025-04: Self-supervised paradigm innovation โ CheXWorld brings JEPA world-modeling to radiographs; multi-teacher distillation agglomerates heavy FMs into lightweight students.
- 2025-05: Scaling laws and generalist-VLM adaptation โ BioVFM-21M characterizes SSL scaling; MedBridge repurposes frozen general VLMs via MoE; XMedGPT adds grounding/uncertainty/prognosis.
- 2025-06โ07: Multimodal multi-disease vision FMs (MerMED-FM) and a turn toward trustworthiness โ de-identification (DCM-DeID), latent-fragility auditing (LAPD).
- 2025-08โ09: Benchmarking-as-contribution and continual learning consolidate โ generalist-vs-specialist studies (ocular), data scaling laws for radiology, UNICON/EWC-diffusion-replay, and Citrus-V unified grounding.
- 2025-10: External-validity reckoning โ CardioBench and ProstNFound+ (first prospective FM validation) test whether FMs survive "beyond the lab"; video models probed as zero-shot medical learners.
- 2025-11: Health-system-scale 3D FMs (NeuroVFM/Vol-JEPA, 5.24M volumes) and unified grounded VLMs (UMind-VL) define the frontier; comparative encoder benchmarks proliferate.
- 2025-12: Generalist 3D-via-2D adaptation (AnyMC3D), retrieval-guided continual learning (PRIMED), surgical FMs (LapFM), and FMs used as semantic priors inside other pipelines (PGMP).
๐ Trend¶
Where the field stands
Foundation models in medicine matured in 2025 from a collection of single-modality encoders into a contested design space spanning vision-only self-supervised backbones, medical vision-language models (VLMs), and unified generalist systems with pixel-level grounding. The year's central tension is specialist vs. generalist: domain-specific medical pretraining (RETFound, MedSAM, CT-FM) still wins on data efficiency and fine-grained tasks, but rapidly scaling general-purpose encoders (DINOv2/v3, SigLIP2, video LVMs) increasingly close the gap, and several benchmarks now show no single model dominates. Three forces shape almost every paper: the impossibility of centralizing patient data (driving federated, model-merging, and continual-learning approaches), the scarcity of labeled volumetric/3D data (driving SSL paradigms like JEPA and SimCLR), and the demand for clinical trust (driving grounding, uncertainty, robustness, and prospective validation). By year's end the frontier had shifted toward 3D/volumetric models trained on uncurated health-system data, unified perception-plus-reasoning VLMs, and rigorous benchmarking that interrogates how FMs are adapted and evaluated rather than chasing single-task SOTA.
โก๏ธ Read the full trend analysis
๐ Papers (43)¶
โก๏ธ Paper list โ 43 papers, grouped by month. ยท โญ 14 key papers (see Key papers in the left nav).