PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 26, 2026PLoS ONE0 citationsOpen Access

CMAP-Fusion: A cross-modal feature selection and model pruning framework for laboratory and imaging data

View Full Paper
CLChong LiuLYLei YangJLJinmeng Lei

Key Points

  • The aim is to develop a framework that optimizes cross-modal fusion of medical imaging and laboratory data for better disease diagnosis.
  • Proposed CMAP-Fusion model uses encoding alignment, redundant pruning, and fusion prediction.
  • Utilized ViT-B/16 for imaging feature extraction and dimension alignment.
  • Employed the SmartTrim dynamic pruning module to reduce redundancy and focus on key features.
  • Achieved accuracies of 95.3% for COVID-19 Radiography, 89.7% for ISIC Skin Cancer, and 93.6% for ChestX-ray14, improving by 3.1% to 4.1% over baselines.
  • Reduced model parameters by 44.2% and computational complexity by over 43%.
  • Significantly enhanced cross-modal similarity and feature sparsity compared to existing methods.

Abstract

Cross-modal fusion of medical imaging and laboratory data is a key pathway for accurate diagnosis of diseases, yet it is constrained by issues such as the modal heterogeneity gap, accumulation of feature redundancy, and efficiency imbalance. Existing methods struggle to balance precision and clinical adaptability, and some rely on simulated data leading to limited generalization ability. To address these challenges, we propose the Cross-Modal Alignment-Pruning Fusion model (CMAP-Fusion), which achieves optimization through modular collaboration of “encoding alignment → redundant pruning → fusion prediction”: ViT-B/16 is used to complete imaging feature extraction and dimension alignment, the SmartTrim dynamic pruning module screens key features and reduces redundancy, and the Cross-Modal Transformer (CMT) mines deep associations between dual modalities. Experiments on the COVID-19 Radiography Dataset, ISIC Skin Cancer Dataset, and ChestX-ray14 Dataset demonstrate that the model achieves accuracies of 95.3%, 89.7%, and 93.6% respectively, representing an improvement of 3.1% to 4.1% compared with optimal baselines. Meanwhile, the number of parameters is reduced by 44.2%, computational complexity is decreased by more than 43%, and cross-modal similarity and feature sparsity are significantly superior to baselines. This model realizes the synergistic optimization of “precision-efficiency-generalization,” providing an efficient solution for medical cross-modal fusion. In the future, we will expand to multi-source modalities and multi-disease scenarios, strengthen clinical multi-center validation, further improve the model’s interpretability and clinical acceptance, and facilitate the lightweight deployment of medical AI.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Liu et al. (2026) studied this question.

synapsesocial.com/papers/69edadba4a46254e215b552bhttps://doi.org/10.1371/journal.pone.0346875
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1CMJRT: Cross-Modal Joint Representation Transformer for Multimodal Sentiment Analysis2022 · 24 citations
  2. 2Analysis of the ISIC image datasets: Usage, benchmarks and recommendations2021 · 301 citations
  3. 3Classification of Paediatric Pneumonia Using Modified DenseNet-121 Deep-Learning Model2024 · 97 citations
  4. 4HFT-Net: Hybrid Fusion Transformer Network for Multi-Source Breast Cancer Classification2025 · 6 citations
  5. 5DIMF-Nets: depth-informed cross-modal fusion in three-stream networks for enhanced unsupervised video object segmentation2024 · 4 citations