GCD (Medical Multi-Modal Generalized Category Discovery), which exploits image- text pairs to jointly recognize known classes and discover novel categories in medical images. To address the varying contribution of different modalities across samples, we develop a Dynamic Expert Fusion module to automatically learn sample-specific modality weights, and further design a Local Experts Balancing mechanism to preserve the discriminative power of individual modalities. By integrating global and local perspectives, our framework adaptively balances modality contributions and enhances multi-modal robustness. Subsequently, to enable the discovery of novel unknown categories during training, we propose a Category Diffusion module grounded in the Metropolis- Hastings framework. This module adaptively merges and splits categories, allowing the model to simultaneously recognize known classes and uncover previously unseen categories during training, without requiring any prior knowledge about the unknown categories. Extensive experiments on two public multi-modal datasets (MIMIC-CXR and PatchGastric), together with a private multi-modal fundus dataset, MM-Retina, demonstrate that our method consistently improves clustering performance on both known and unknown categories compared with existing approaches.
Feng et al. (2026) studied this question.