ABSTRACT Deepfake detection models achieve high accuracy, yet their interpretability remains underexplored. This study presents a unified evaluation pipeline for post hoc visual explanations, grounded in the Co‐12 attributes of explanation quality and operationalised through a structured framework that assesses Coherence, Composition, Correctness, Completeness, Compactness, Covariate completeness and Continuity. Applying this pipeline to 16 representative explanation methods reveals systematic differences across methodological categories. CAM‐based approaches demonstrate strong spatial coherence and temporal continuity; gradient‐based techniques such as Guided Backprop and LRP yield compact and accurate attributions; and redistribution‐based methods including ExcitationBP and Deep Taylor maintain high consistency across evaluation conditions. In contrast, perturbation‐based approaches such as SHAP and LIME exhibit weaker localisation and reduced temporal stability. By enabling controlled, attribute‐level comparison of explanation methods, the proposed pipeline bridges conceptual interpretability frameworks and empirical analysis, offering practical guidance for the development and deployment of interpretable deepfake detectors in forensic and auditing applications. The source code is publicly available at: https://github.com/junxinchenieee/EAI‐Deepfake‐Detection .
Li et al. (Wed,) studied this question.