PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 2, 2026Journal of Imaging0 citationsOpen Access

AACNN-ViT: Adaptive Attention-Augmented Convolutional and Vision Transformer Fusion for Lung Cancer Detection

View Full Paper
MRMohammad Ishtiaque RahmanARAmrina Rahman

Key Points

  • The study aims to enhance lung cancer detection through the integration of convolutional and transformer models.
  • Developed AACNN-ViT hybrid framework combining CNN and ViT models.
  • Implemented an adaptive attention-based fusion mechanism for representation learning.
  • Used a hybrid loss function integrating focal loss and categorical cross-entropy.
  • Performed evaluations on the IQ-OTH/NCCD dataset comparing multiple model architectures.
  • AACNN-ViT achieved 96.97% accuracy on the validation set.
  • Macro-averaged precision, recall, and F1 scores were 0.9588, 0.9352, and 0.9458 respectively.
  • Substantial improvement in minority-class recognition with benign recall at 0.8333.
  • Micro-average AUC of 0.992 from one-vs.-rest ROC analysis indicates strong class separability.

Abstract

Lung cancer remains a leading cause of cancer-related mortality. Although reliable multiclass classification of lung lesions from CT imaging is essential for early diagnosis, it remains challenging due to subtle inter-class differences, limited sample sizes, and class imbalance. We propose an Adaptive Attention-Augmented Convolutional Neural Network with Vision Transformer (AACNN-ViT), a hybrid framework that integrates local convolutional representations with global transformer embeddings through an adaptive attention-based fusion module. The CNN branch captures fine-grained spatial patterns, the ViT branch encodes long-range contextual dependencies, and the adaptive fusion mechanism learns to weight cross-representation interactions to improve discriminability. To reduce the impact of imbalance, a hybrid objective that combines focal loss with categorical cross-entropy is incorporated during training. Experiments on the IQ-OTH/NCCD dataset (benign, malignant, and normal) show consistent performance progression in an ablation-style evaluation: CNN-only, ViT-only, CNN-ViT concatenation, and AACNN-ViT. The proposed AACNN-ViT achieved 96.97% accuracy on the validation set with macro-averaged precision/recall/F1 of 0.9588/0.9352/0.9458 and weighted F1 of 0.9693, substantially improving minority-class recognition (Benign recall 0.8333) compared with CNN-ViT (accuracy 89.09%, macro-F1 0.7680). One-vs.-rest ROC analysis further indicates strong separability across all classes (micro-average AUC 0.992). These results suggest that adaptive attention-based fusion offers a robust and clinically relevant approach for computer-aided lung cancer screening and decision support.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Rahman et al. (2026) studied this question.

synapsesocial.com/papers/6980fde8c1c9540dea80fa19https://doi.org/10.3390/jimaging12020062
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Evaluation metrics and statistical tests for machine learning2024 · 1,218 citations
  2. 2Ensemble approach of transfer learning and vision transformer leveraging explainable AI for disease diagnosis: An advancement towards smart healthcare 5.02024 · 28 citations
  3. 3Explaining explainability: The role of XAI in medical imaging2024 · 30 citations
  4. 4Method for Diagnosis of Acute Lymphoblastic Leukemia Based on ViT‐CNN Ensemble Model2021 · 123 citations
  5. 5Visual Attention-Driven Hyperspectral Image Classification2019 · 251 citations