Key points are not available for this paper at this time.
Intracranial hemorrhage (ICH) is a common emergency worldwide. We systematically evaluated studies on ICH detection and subtype classification using deep learning (DL) models. The study protocol was registered with PROSPERO (CRD420251081485). Studies published between 2020 and June 2025 were searched in PubMed, Scopus, Web of Science, IEEE Xplore, and ScienceDirect. Risk of bias was assessed using QUADAS-2. Following screening and selection, 90 studies were included in the qualitative assessment. Among these, studies reporting or allowing derivation of TP, FP, TN, and FN were included in the meta-analysis. Across 46 ICH detection studies, pooled estimates differed by analysis level. In scan-level evaluations, pooled sensitivity was 90% (95% CI: 88%-92%) and pooled specificity was 94% (95% CI: 92%-96%) (SROC-AUC 0.963). In slice-level evaluations, pooled sensitivity was 95% (95% CI: 93%-97%) and pooled specificity was 97% (95% CI: 91%-99%) (SROC-AUC 0.984). Because slice-level and scan-level tasks are not directly comparable and heterogeneity was substantial, we emphasized scan-level performance for clinical interpretation and reported prespecified subgroup meta-analyses to contextualize variability. Across eleven subtype classification studies, pooled sensitivities ranged from 78% to 89% across hemorrhage subtypes, with pooled specificities of 96%-99%, indicating clinically relevant variability across subtypes. Given that false negatives in time-critical subtypes can delay urgent management, the current evidence supports these systems primarily as human-in-the-loop decision support rather than stand-alone triage.
Habek et al. (2026) studied this question.