Diagnostic Accuracy of Three Large Language Models on Clinical Images Varies by Skin Tone: Findings From the Stanford Diverse Dermatology Images Dataset | Synapse
May 7, 2026International Journal of Dermatology0 citations
Diagnostic Accuracy of Three Large Language Models on Clinical Images Varies by Skin Tone: Findings From the Stanford Diverse Dermatology Images Dataset
This research aims to evaluate how accurately three large language models diagnose clinical images across different skin tones.
Used the Stanford Diverse Dermatology Images dataset to assess the models.
Analyzed diagnostic accuracy across various skin tones.
Data is openly available for further examination.
Diagnostic performance varied significantly by skin tone, with some tones receiving less accuracy.
The findings highlight a potential bias in model predictions.
Implications suggest a need for more diverse training data in AI.
Abstract
The data that support the findings of this study are openly available in Stanford's Diverse Dermatology Image at https://aimi.stanford.edu/datasets/ddi-diverse-dermatology-images.