PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 31, 2026Scientific Reports0 citationsOpen Access

MSML-DenseXmer: harnessing vision transformers through integration with novel dense networks for medical image fusion

DPDisha Mohini PathakDJDhyanendra JainSSSomya Srivastava

Key Points

  • This study aims to improve the quality of medical image fusion by using a novel deep learning framework that integrates DenseNet and Swin Transformer architectures.
  • Developed a deep-learning framework combining DenseNet for local features and Swin Transformer for global details.
  • Utilized L1, L2, and infinity norm for generating attention weights in feature fusion.
  • Trained on Whole Brain Atlas and Lung-PET-CT-Dx datasets with a unique loss function.
  • Achieved over 9.94% improvement in MRI-SPECT fusion quality.
  • Observed above 6.82% increase in MRI-PET fusion effectiveness.
  • Reported a minimum of 18.37% enhancement in MRI-CT fusion quality.

Abstract

Abstract The integration of multiple modalities in medical imaging allows a thorough representation of structural and functional details, resulting in improved diagnosis and treatment. Deep learning methods outperform conventional methods by automating the extraction of pertinent features and fusing them while preserving both structural and textural integrity. Existing methods lack the ability to capture complex global structures, small-scale textural features, and long-range dependencies, which causes incomplete feature representation. This study presents a novel deep-learning framework that combines an improved DenseNet for capturing local fine- grained features with a Swin Transformer for extracting global structural details and long- range relationships, thereby facilitating a more comprehensive fused output. A modified hybrid approach utilizing L1, L2 and infinity norm is used to generate attention weights in the feature fusion-oriented row-column vector dimension technique. The model is trained on different modalities in the Whole Brain Atlas dataset and the Lung-PET-CT-Dx dataset using a novel loss function. This function improves fusion by integrating pixel loss, structural similarity, and textural preservation. The evaluation of the fused image’s quality involves multiple metrics that assess image clarity, structural integrity, feature retention, contrast improvement, and overall visual accuracy, providing a thorough analysis. The MSML-DenseXmer framework demonstrates improved performance compared to existing approaches across multiple medical imaging modalities. Specifically, it achieves over 9.94% rise in MRI-SPECT fusion, above 6.82% gain in MRI-PET fusion, a minimum of 18.37% increase in MRI-CT fusion, and at least 2.83% gain on the Lungs PET-CT dataset, indicating its potential in improving fusion quality.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Pathak et al. (2026) studied this question.

synapsesocial.com/papers/6a1bd2515783ba022b6fdbd7https://doi.org/10.1038/s41598-026-52181-8
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Enhanced Deep Learning Techniques for Multimodal Medical Image Fusion2024 · 2 citations
  2. 2DBCFuse: Dual-Branch Cross-Scale Feature Decomposition for MRI-PET Fusion2026
  3. 3Multimodal Medical Image Fusion Method based on the Swin Transformer and Self-supervised Contrast Learning2024 · 1 citations
  4. 4MSFusion: Multi-Scale Cross-Modal Fusion with Adaptive Attention for Multimodal Medical Image Fusion2026
  5. 5Multimodal medical image fusion and classification using deep learning techniques2024 · 12 citations