Skip to content

Imaging-anchored Multiomics in Cardiovascular Disease: Integrating Cardiac Imaging, Bulk, Single-cell, and Spatial Transcriptomics

🕒 Published (v1): 2026-01-10 23:30 UTC · Source: Arxiv · link

Ask a follow-up

Open an assistant pre-loaded with this paper's context.

💬 Ask ChatGPT✦ Ask Claude

TL;DR

This review proposes an imaging-anchored multiomics framework for cardiovascular disease in which cardiac MRI, CT, and echocardiography serve as the spatial reference frame, while bulk, single-cell, and spatial transcriptomics provide cell-type- and location-specific molecular context. The central goal is to build joint computational representations that allow imaging phenotypes (e.g., late gadolinium enhancement, plaque morphology) to be decoded into underlying molecular states. Foundation models for single-cell omics, spatial transcriptomics, and multimodal medical imaging are identified as key enablers of this vision.

Problem

Cardiac imaging and transcriptomic data (bulk, scRNA-seq, spatial) are generated routinely in clinical and research cohorts but analyzed in entirely separate pipelines, preventing mechanistic interpretation of imaging phenotypes in terms of cell states, pathways, and spatial niches. Existing reviews treat imaging and molecular layers as parallel streams rather than constructing joint, cross-scale representations.

Method

The review taxonomizes and synthesizes representation-learning and fusion strategies across modalities: - Imaging encoders: CNNs/3D-CNNs for segmentation and function; Vision Transformers and masked autoencoders (MAEs) for self-supervised pre-training on unlabeled cardiac archives, producing global/regional/voxel-level embeddings. - Omics encoders: PCA/ICA and VAEs for bulk RNA-seq; scVI-style hierarchical VAEs for single-cell count data; MOFA+ for cross-omic factor decomposition; emerging large-scale single-cell/spatial foundation models (transformers, state-space models on tens of millions of cells). - Spatial transcriptomics encoders: Graph convolutional networks (SpaGCN, BayesSpace) over spot graphs; multimodal encoders jointly processing H&E patches and expression vectors; cross-attention/diffusion architectures trained in a foundation-model regime. - Fusion architectures: Early (feature concatenation), intermediate (DCCA, multimodal VAEs with product-of-experts/mixture-of-experts, cross-modal autoencoders), and late (ensemble of unimodal models) fusion, plus hybrid designs. Graph-based fusion uses GNNs over feature graphs (imaging regions ↔ pathway/gene nodes) and patient-similarity graphs. Contrastive objectives (CLIP-style) align imaging–omics pairs in shared latent spaces. - Integrative pipelines: Radiogenomics (correlating imaging features with gene/pathway scores, causal Mendelian randomization); spatial molecular alignment (multi-stage histology-to-in-vivo registration with deformable warping); image-based gene-expression prediction ("virtual transcriptomics") predicting pathway scores from imaging features.

Key Contributions

  • Introduces the imaging-anchored multiomics perspective as a unifying framework, positioning cardiac imaging as the primary spatial coordinate system for cross-scale integration.
  • Comprehensive taxonomy of fusion strategies (early/intermediate/late/graph/contrastive) with explicit analysis of trade-offs for missing-data handling and small paired sample sizes.
  • Systematic review of public multimodal cardiovascular datasets (UK Biobank CMR+genomics ~100k, MESA, HCMR, MIMIC-IV-ECG, EchoNet, Cardiac Atlas Project, spatial MI atlas, coronary plaque spatial omics).
  • Surveys representative integrative pipelines: radiotranscriptomic perivascular signatures (CT radiomics → vascular inflammatory gene-expression scores), cross-modal ECG–CMR autoencoders at biobank scale, spatial multi-omic maps of human MI and atherosclerotic plaque.
  • Identifies single-cell/spatial foundation models and multimodal medical foundation models as the convergent technology enabling large-scale imaging-anchored multiomics translation.
  • Discusses practical failure modes: batch effects, domain shift across scanners/protocols, overfitting in small paired cohorts, tissue deformation artifacts in spatial alignment.

Results

The paper is a review; no original experimental results or new benchmark numbers are reported. Key empirical findings cited from the literature: - Cross-modal cardiovascular autoencoder (ECG + CMR + clinical traits): demonstrated scalable genotype association and CMR imputation from ECG at tens-of-thousands scale. - Radiotranscriptomic perivascular score (coronary CTA radiomics mapped to perivascular adipose gene-expression): predicts cardiac mortality and MI beyond standard CT and clinical risk scores. - Spatial multi-omic map of human MI (31 samples, 23 patients; snRNA-seq + snATAC-seq + Visium): delineates infarct core, border zone, and remote myocardium by cell state and pathway. - Coronary plaque spatial transcriptomics (multiple cohorts): spatially resolves immune niches and gene-expression modules co-localizing with high-risk CT plaque morphologies (low-attenuation plaque, thin-cap fibroatheroma).

Limitations

  • No new methods or benchmarks are introduced; synthesis relies on heterogeneous literature without head-to-head comparisons.
  • Paired imaging–omics datasets remain small (tens to hundreds of patients), making deep fusion model training prone to overfitting and site confounding.
  • Spatial alignment across scales (MRI voxel → histology → spatial spot) introduces cumulative registration uncertainty from tissue deformation and sectioning artifacts; no standardized protocol exists.
  • Most spatial foundation model demonstrations are in oncology; cardiac/vascular validation is prospective rather than demonstrated.
  • Benchmarking frameworks with standardized splits and external validation for imaging–omics integration are largely absent.
  • Batch effects from scanner heterogeneity, imaging protocol variation, and omics platform differences are partially addressed but remain a dominant confounder.
  • Scope is limited to transcriptomics; other omics (epigenomics, metabolomics, proteomics) are acknowledged but not deeply integrated into the review's framework.

Relevance to Foundation Models in Medicine

This review is directly relevant to medical foundation models as it frames single-cell/spatial foundation models (large transformers and state-space models trained on tens of millions of cells across tissues) and multimodal medical foundation models (coupling imaging, text, and clinical data) as the key architectural convergence points for imaging-anchored cardiovascular multiomics. It explicitly positions these models as transferable encoders that can provide generic representations for downstream cross-modal fusion, while also surfacing critical challenges—domain-specific fine-tuning, curation, and auditing—that must be addressed before such backbones generalize to cardiovascular multiomics. For researchers tracking foundation models in medicine, this paper maps the current gap between general-purpose medical FMs and the specialized registration, spatial alignment, and missing-data handling required for cross-scale cardiovascular integration.