Accurately predicting individual responses to Anti-Vascular Endothelial Growth Factor (Anti-VEGF) efficacy in diabetic macular edema (DME) remains a critical challenge in personalized ophthalmic care. Existing methods often rely on unimodal data or suffer from ineffective multimodal feature extraction and fusion, leading to modality redundancy and performance degradation. To address these limitations, we propose MVTT-GMamba, a novel multimodal learning framework that integrates optical coherence tomography (OCT) images and structured clinical indicators for early and accurate Anti-VEGF efficacy prediction. At the core of MVTT-GMamba is a feature-wise heterogeneous graph reasoning paradigm that explicitly models inter-patient and inter-feature relations, together with an adaptive, graph-guided prediction head that progressively anneals structural priors into the classifier. Building on this core, we adopt domain-tailored MambaVision and TabTransformer encoders and an early cross-attention fusion module to realize fine-grained multimodal representation learning. Extensive experiments on both a private clinical dataset (DMETHERA-ECSAHZU) and the public APTOS2021 benchmark demonstrate that MVTT-GMamba consistently outperforms state-of-the-art methods across all evaluation metrics. In addition, Grad-CAM visualizations reveal that the model attends to clinically relevant retinal regions, providing enhanced interpretability. Code is available at: https://github.com/DME666/DME.
Wu et al. (Thu,) studied this question.