PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 16, 2026Journal of Mechanics in Medicine and Biology0 citations

Generalization Performance of Deep Learning Models for Skin Lesion Classification

View Full Paper
WCWen-Yen ChangCKChing-Chang KuoJXJia-Lang Xu

Key Points

  • This research aims to evaluate the generalization performance of various deep learning models on skin lesion classification across different datasets.
  • Benchmarking seven deep learning architectures including ShuffleNetV2, RegNetY, and MobileNetV3.
  • Utilizing three publicly available datasets for training, validation, and testing.
  • Employing five metrics: Accuracy, Precision, Sensitivity, Specificity, and F1-score for evaluation.
  • Models achieved over 99% accuracy on the training dataset.
  • Validation accuracy declined to 76-84%, with MobileNetV3 leading at 83.58%.
  • Testing accuracies ranged from 64-69%, with RegNetY at 68.89% and MobileNetV3 at 68.62%.
  • Specificity remained high, above 94%, indicating reliability in identifying negative cases.

Abstract

The classification of skin lesions is essential for early diagnosis and improved clinical outcomes in dermatology. Although deep learning techniques have significantly advanced medical image analysis, ensuring that these models can generalize effectively across diverse datasets remains a critical requirement for real-world applications. This study benchmarked seven state-of-the-art deep learning architectures, including ShuffleNetV2, RegNetY, MobileNetV3, MobileViT, EfficientNetB0, MnasNet, and DeiT. This study used three publicly available datasets to employed and model performance. It was evaluated across training, validation, and testing datasets using five standard metrics: Accuracy, Precision, Sensitivity, Specificity, and F1-score. All models achieved near-perfect performance on the training dataset, exceeding 99% across all metrics and demonstrating strong learning capacity. On the validation dataset, performance declined to a range of 76 to 84% accuracy, with MobileNetV3 achieving the best results at 83.58%, followed by RegNetY at 82.34 %. On the independent testing dataset, accuracies further decreased to between 64 and 69 %, with RegNetY reaching 68.89% and MobileNetV3 reaching 68.62%, outperforming the other models. Specificity remained consistently high, above 94 percent across all models, indicating reliable identification of negative cases. The findings highlight the importance of rigorous evaluation across multiple datasets to assess true model generalization. While all tested models demonstrated strong learning ability, MobileNetV3 and RegNetY consistently showed superior robustness, suggesting their potential suitability for clinical deployment.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Chang et al. (2026) studied this question.

synapsesocial.com/papers/69e07e582f7e8953b7cbf676https://doi.org/10.1142/s0219519426500259
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Classification of Skin Diseases with Different Deep Learning Models and Comparison of the Performances of the Models2024
  2. 2Ensemble Method of Pre-Trained Models for Classification of Skin Lesion Images2025 · 2 citations
  3. 3Enhancing Dermatological Diagnostics with EfficientNet: A Deep Learning Approach2024 · 5 citations
  4. 4Accurate Deep Learning Algorithms for Skin Lesion Classification2024
  5. 5Bridging the Gap Between Theoretical Performance and Clinical Utility in Multi-Class Skin Lesion Diagnosis2026 · 2 citations