Skip to content

Foundation Models in Medicine โ€” 2025

๐Ÿ—“๏ธ 2025 in Review

๐Ÿ•’ Timeline

  • 2025-01: Native-3D vision-only CT encoders arrive โ€” CT-FM (148k scans, intra-sample contrastive) and MEDFORM (CT+clinical-tabular contrastive) establish vision-centric, non-text 3D pretraining.
  • 2025-02: Privacy-driven composition emerges โ€” MedForge (open-source-style LoRA merging), Med-LEGO (training-free model fusion), MedTok (graph-aware EHR tokenizer), and RELICT (memorization/replica auditing).
  • 2025-03: Generalist-via-adaptation peaks โ€” continual-learning segmentation (CL-Net, 235 anatomies), unified ReID (MaMI), few-shot domain adaptation (MFM-DA), and the CUPCase finding that GPT-4o beats clinical LLMs on rare cases.
  • 2025-04: Self-supervised paradigm innovation โ€” CheXWorld brings JEPA world-modeling to radiographs; multi-teacher distillation agglomerates heavy FMs into lightweight students.
  • 2025-05: Scaling laws and generalist-VLM adaptation โ€” BioVFM-21M characterizes SSL scaling; MedBridge repurposes frozen general VLMs via MoE; XMedGPT adds grounding/uncertainty/prognosis.
  • 2025-06โ€“07: Multimodal multi-disease vision FMs (MerMED-FM) and a turn toward trustworthiness โ€” de-identification (DCM-DeID), latent-fragility auditing (LAPD).
  • 2025-08โ€“09: Benchmarking-as-contribution and continual learning consolidate โ€” generalist-vs-specialist studies (ocular), data scaling laws for radiology, UNICON/EWC-diffusion-replay, and Citrus-V unified grounding.
  • 2025-10: External-validity reckoning โ€” CardioBench and ProstNFound+ (first prospective FM validation) test whether FMs survive "beyond the lab"; video models probed as zero-shot medical learners.
  • 2025-11: Health-system-scale 3D FMs (NeuroVFM/Vol-JEPA, 5.24M volumes) and unified grounded VLMs (UMind-VL) define the frontier; comparative encoder benchmarks proliferate.
  • 2025-12: Generalist 3D-via-2D adaptation (AnyMC3D), retrieval-guided continual learning (PRIMED), surgical FMs (LapFM), and FMs used as semantic priors inside other pipelines (PGMP).

๐Ÿ“ˆ Trend

Where the field stands

Foundation models in medicine matured in 2025 from a collection of single-modality encoders into a contested design space spanning vision-only self-supervised backbones, medical vision-language models (VLMs), and unified generalist systems with pixel-level grounding. The year's central tension is specialist vs. generalist: domain-specific medical pretraining (RETFound, MedSAM, CT-FM) still wins on data efficiency and fine-grained tasks, but rapidly scaling general-purpose encoders (DINOv2/v3, SigLIP2, video LVMs) increasingly close the gap, and several benchmarks now show no single model dominates. Three forces shape almost every paper: the impossibility of centralizing patient data (driving federated, model-merging, and continual-learning approaches), the scarcity of labeled volumetric/3D data (driving SSL paradigms like JEPA and SimCLR), and the demand for clinical trust (driving grounding, uncertainty, robustness, and prospective validation). By year's end the frontier had shifted toward 3D/volumetric models trained on uncurated health-system data, unified perception-plus-reasoning VLMs, and rigorous benchmarking that interrogates how FMs are adapted and evaluated rather than chasing single-task SOTA.

โžก๏ธ Read the full trend analysis

๐Ÿ“„ Papers (43)

โžก๏ธ Paper list โ€” 43 papers, grouped by month. ยท โญ 14 key papers (see Key papers in the left nav).