PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 6, 2026IEEE Transactions on Medical Imaging1 citations

Disentangled Multi-modal Learning of Histology and Transcriptomics for Cancer Characterization

View Full Paper
YZYupei ZhangXWXiaofei WangALA. M. M. Liu

Key Points

  • The aim is to develop a framework that enhances cancer characterization by addressing challenges in multi-modal learning between histology and transcriptomics.
  • Developed a disentangled multi-modal fusion module to reduce heterogeneity.
  • Introduced a gene expression consistency strategy for aligning transcriptomic signals at various WSI magnifications.
  • Created a knowledge distillation strategy to allow for WSI-only model inference without relying on paired data.
  • Implemented a token aggregation module to enhance inference efficiency by reducing redundancy.
  • Demonstrated improved cancer diagnosis and prognosis metrics compared to existing methods.
  • Achieved higher accuracy in survival prediction across diverse clinical scenarios.

Abstract

Histopathology remains the gold standard for cancer diagnosis and prognosis. With the advent of transcriptome profiling, multi-modal learning combining transcriptomics with histology offers more comprehensive information. However, existing multi-modal approaches are challenged by intrinsic multi-modal heterogeneity, insufficient multi-scale integration, and reliance on paired data, restricting clinical applicability. To address these challenges, we propose a disentangled multi-modal framework with four contributions: 1) To mitigate multimodal heterogeneity, we decompose WSIs and transcriptomes into tumor and microenvironment subspaces using a disentangled multi-modal fusion module, and introduce a confidence-guided gradient coordination strategy to balance subspace optimization. 2) To enhance multi-scale integration, we propose an inter-magnification geneexpression consistency strategy that aligns transcriptomic signals across WSI magnifications. 3) To reduce dependency on paired data, we propose a subspace knowledge distillation strategy enabling transcriptome-agnostic inference through a WSI-only student model. 4) To improve inference efficiency, we propose an informative token aggregation module that suppresses WSI redundancy while preserving subspace semantics. Extensive experiments on cancer diagnosis, prognosis, and survival prediction demonstrate our superiority over state-of-the-art methods across multiple settings. Code is available at GitHub.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhang et al. (2026) studied this question.

synapsesocial.com/papers/69aa6f3c531e4c4a9ff59474https://doi.org/10.1109/tmi.2026.3669968
Ask AI
Helpful
Bookmark
Share
View Full Paper