Multimodal Medical Image Fusion (MMIF) can integrate complementary information from different imaging modalities to generate more comprehensive fused images and improve the reliability of clinical diagnosis. However, it is difficult to handle the subtle local details during feature extraction and cross-modal interaction, which may lead to the loss of critical diagnostic information. In this paper, we propose neighborhood-attention-based Multiscale Alignment and hierarchical Reconstruction for multimodal medical image Fusion (MARFusion). First, we present a Neighborhood Attention Network (NAnet) that dynamically allocates pixel-level attention within local neighborhoods to enhance fine-grained detail preservation. Next, we employ NAnet to extract multiscale feature maps and randomly crop them to obtain patch-level representations. Then, we apply patch-level contrastive learning at each scale to promote cross-modal alignment in local regions. Finally, we adopt a hierarchical image reconstruction strategy for information retention, i.e., multiscale patch-level local image reconstruction to emphasize subtle cues and cross-scale global image reconstruction to maintain semantic consistency. Extensive experiments validate the effectiveness of MARFusion and demonstrate the competitive performance compared with state-of-the-art methods on PET–MRI and SPECT–MRI medical image fusion tasks.
Cao et al. (Mon,) studied this question.