Artificial intelligence (AI) systems for radiographic caries detection are commonly evaluated using a small set of performance metrics, yet these measures are frequently reported and interpreted without the contextual information required for clinically meaningful conclusions. This narrative review focuses on performance metrics - confusion-matrix measures (sensitivity, specificity, accuracy, predictive values, F1 score and related summaries) and discrimination metrics (ROC-AUC and PR-AUC) -and discusses how their interpretation depends on explicit reporting of the disease definition (grading cut-offs), unit of analysis (tooth surface/patient), decision threshold(s), and prevalence/case-mix. We additionally cover complementary metric-based assessments, including calibration (calibration curves, expected calibration error, Brier score), decision-analytic metrics (decision curve analysis and cost-sensitive clinical loss), and robustness summaries (uncertainty and subgroup/worst-group performance). A structured summary is provided of formulas, interpretive limitations, and the minimum information that should accompany each metric to support transparent reporting, comparability, and transportability.
Ganss et al. (Thu,) studied this question.