The classification of skin lesions is essential for early diagnosis and improved clinical outcomes in dermatology. Although deep learning techniques have significantly advanced medical image analysis, ensuring that these models can generalize effectively across diverse datasets remains a critical requirement for real-world applications. This study benchmarked seven state-of-the-art deep learning architectures, including ShuffleNetV2, RegNetY, MobileNetV3, MobileViT, EfficientNetB0, MnasNet, and DeiT. This study used three publicly available datasets to employed and model performance. It was evaluated across training, validation, and testing datasets using five standard metrics: Accuracy, Precision, Sensitivity, Specificity, and F1-score. All models achieved near-perfect performance on the training dataset, exceeding 99% across all metrics and demonstrating strong learning capacity. On the validation dataset, performance declined to a range of 76 to 84% accuracy, with MobileNetV3 achieving the best results at 83.58%, followed by RegNetY at 82.34 %. On the independent testing dataset, accuracies further decreased to between 64 and 69 %, with RegNetY reaching 68.89% and MobileNetV3 reaching 68.62%, outperforming the other models. Specificity remained consistently high, above 94 percent across all models, indicating reliable identification of negative cases. The findings highlight the importance of rigorous evaluation across multiple datasets to assess true model generalization. While all tested models demonstrated strong learning ability, MobileNetV3 and RegNetY consistently showed superior robustness, suggesting their potential suitability for clinical deployment.
Chang et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: