What question did this study set out to answer?

The study aims to create a framework for assessing explainable AI heatmaps in retinal disease classification to ensure clinical applicability.

May 20, 2026Open Access

Clinician-Centered Evaluation Framework for Explainable AI Heatmaps in OCT-Based Retinal Disease Classification

Puntos clave

The study aims to create a framework for assessing explainable AI heatmaps in retinal disease classification to ensure clinical applicability.
Developed a two-phase evaluation framework for explainable AI heatmaps in OCT classification.
Trained a six-class Swin Transformer model using a combined dataset from public and private sources, achieving high accuracy.
Conducted evaluations by specialists using a five-point Likert scale to assess agreement between highlighted regions and model diagnosis.
Achieved 97% accuracy in cross-validation and 91.82% accuracy in external evaluations.
Token contRAST map received the highest specialist ratings for clinical plausibility.
Grad-CAM++ and Cosine-Grad Fusion Map (CGFM) were rated lower, indicating varying effectiveness of XAI methods.

Resumen

This study presents a two-phase framework for selecting clinically plausible explainable artificial intelligence (XAI) heatmaps for retinal optical coherence tomography (OCT) classification. A six-class Swin Transformer model was trained and validated using a combined dataset consisting of a subset of the public OCT-C8 dataset and private data from a Greek tertiary hospital and externally evaluated on an independent dataset from a private ophthalmological institute. Diagnostic performance was high, achieving 97% accuracy in cross-validation and 91.82% on external evaluation. In Phase 1, one ophthalmologist and one artificial intelligence (AI) specialist independently assessed 100 heatmaps per method based on visual quality and anatomical plausibility, reducing the candidate methods to three. In Phase 2, 21 specialists evaluated the selected methods across multiple cases using a five-point Likert scale reflecting agreement between highlighted regions and the model diagnosis. The proposed Token contRAST map (TRAST) achieved the highest ratings, followed by Gradient-weighted Class Activation Mapping (Grad-CAM++), while Cosine-Grad Fusion Map (CGFM) showed the lowest performance. These findings reflect clinical plausibility rather than direct model interpretability and indicate that effective XAI in OCT imaging requires not only technical performance but also structured expert evaluation. The proposed framework provides a practical approach for selecting explanation methods suitable for clinical use in ophthalmology.

Leer artículo completoexternamente

Me gusta

Guardar

Ver artículo completo

Cite This Study

Maliagkani et al. (Sat,) studied this question.

synapsesocial.com/papers/6a0d5000f03e14405aa9b8d2 https://doi.org/https://doi.org/10.3390/jimaging12050211

Me gusta

Guardar

Ver artículo completo