PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 7, 2026IEEE Transactions on Medical Imaging0 citations

Leveraging Image-text Pairs for Generalized Category Discovery in Medical Image Classification

View Full Paper
WFWei FengBWBingjie WangZWZhonghua Wang

Key Points

  • The aim is to enhance medical image classification by discovering both known and novel categories through a multi-modal approach.
  • Developed GCD that combines image-text pairs for classification and category discovery.
  • Implemented Dynamic Expert Fusion to learn sample-specific modality weights.
  • Introduced Local Experts Balancing to maintain modality discrimination.
  • Proposed the Category Diffusion module based on Metropolis-Hastings for category merging and splitting.
  • GCD method improves clustering performance on known and unknown categories.
  • Demonstrated robustness across multiple datasets including MIMIC-CXR and PatchGastric.
  • Consistent enhancement in classification results compared to existing methodologies.

Abstract

GCD (Medical Multi-Modal Generalized Category Discovery), which exploits image- text pairs to jointly recognize known classes and discover novel categories in medical images. To address the varying contribution of different modalities across samples, we develop a Dynamic Expert Fusion module to automatically learn sample-specific modality weights, and further design a Local Experts Balancing mechanism to preserve the discriminative power of individual modalities. By integrating global and local perspectives, our framework adaptively balances modality contributions and enhances multi-modal robustness. Subsequently, to enable the discovery of novel unknown categories during training, we propose a Category Diffusion module grounded in the Metropolis- Hastings framework. This module adaptively merges and splits categories, allowing the model to simultaneously recognize known classes and uncover previously unseen categories during training, without requiring any prior knowledge about the unknown categories. Extensive experiments on two public multi-modal datasets (MIMIC-CXR and PatchGastric), together with a private multi-modal fundus dataset, MM-Retina, demonstrate that our method consistently improves clustering performance on both known and unknown categories compared with existing approaches.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Feng et al. (2026) studied this question.

synapsesocial.com/papers/69fbe2b3164b5133a91a22a7https://doi.org/10.1109/tmi.2026.3689859
Ask AI
Helpful
Bookmark
Share
View Full Paper