Background: Melanoma remains a leading cause of cancer-related mortality, with early detection being the primary determinant of survival. The emergence of MLLMs offers a potential paradigm shift in accessible screening. However, the diagnostic reliability and safety of these general-purpose models in oncology remain insufficiently characterized. Methods: This study performed a head-to-head comparison of GPT-5, Gemini 3, and Grok 4 to evaluate their efficacy as first-level screening tools for cutaneous melanoma. A retrospective analysis was conducted using a balanced dataset of 100 clinical images (50 histopathologically confirmed benign, 50 malignant) from the ISIC archive. Results: Gemini 3 achieved the highest overall accuracy (71%) and specificity (94%), while Grok 4 demonstrated the highest sensitivity (52%). All models exhibited a critical deficit in sensitivity, missing approximately half of the malignant lesions. Statistical testing revealed no significant performance differences between the models (p > 0.05). Notably, Gemini 3 exhibited severe overconfidence, maintaining a high CI (84.62%) even during false-negative predictions, whereas GPT-5 and Grok 4 showed better calibration with a significant drop in confidence upon incorrect diagnosis. Conclusions: While current MLLMs possess a foundational capacity for dermatological analysis, their unacceptably low sensitivity and potential for overconfident misdiagnosis render them unsafe as standalone screening tools. At present, MLLMs should only be utilized as complementary tools under strict clinical supervision.
Andrei et al. (Thu,) studied this question.