The TSDNet model achieved 96.19% video-level accuracy and 86.86% frame-level accuracy for detecting ventricular septal defects in pediatric echocardiography videos on the internal test set.
Observational (n=384)
Sí
Does the TSDNet model accurately detect ventricular septal defects in pediatric echocardiographic videos compared to standard object detection models?
The TSDNet model provides highly accurate and robust automated detection of ventricular septal defects from pediatric echocardiography videos, demonstrating strong potential to enhance diagnostic capabilities in primary healthcare settings.
ABSTRACT This study aimed to develop an automated transthoracic echocardiography (TTE)‐based system for detecting ventricular septal defects (VSD) by jointly analyzing systolic‐phase temporal shunt dynamics and precisely localizing defect regions. This multicenter retrospective study analyzed transthoracic echocardiography videos from 324 pediatric patients (150 VSD ‐positive and 174 VSD ‐negative), supplemented by an external validation cohort of 60 cases. To address challenges inherent to dynamic cardiac hemodynamics and view‐dependent variability, we developed a Temporal–Spatial Decoupled Network ( TSDNet ). The methodological contributions of TSDNet are fourfold: (1) a temporal–spatial decoupling strategy that separates transient systolic shunt dynamics from static anatomical localization, overcoming the limitations of conventional spatiotemporal coupling; (2) a dual‐branch architecture that integrates a UniFormer ‐based temporal classifier with a YOLOv5 ‐based spatial detector to jointly perform video‐level diagnosis and frame‐level shunt localization; (3) multi‐view fusion across five standard echocardiographic views to capture complementary anatomical structures and hemodynamic signatures; and (4) a comprehensive evaluation framework encompassing accuracy, recall, specificity, F1 ‐score, PPV , and NPV , supplemented by experiments across multiple temporal window lengths and comparative analyses against state‐of‐the‐art object detection models. T SDNet demonstrated high performance in both video‐level classification and frame‐level VSD detection. Video‐level accuracy on the internal test set reached 96.19%, with external performance remaining consistently strong across varying temporal windows (87.37%–90.37% accuracy). For frame‐level detection, TSDNet achieved 86.86% accuracy on the internal test set and 73.75% on the external cohort, with longer temporal windows yielding progressively improved results. Across all evaluated metrics, TSDNet outperformed leading detection frameworks—including Sparse R‐ CNN , DINO , FCOS , RetinaNet , Faster R‐ CNN , and DCNv2 . Multi‐view analysis further confirmed stable detection performance, particularly in the A4C , PSAX , and subRVOT views, underscoring the effectiveness of combining diverse echocardiographic perspectives with temporal–spatial decoupling for robust VSD identification. T SDNet provides accurate, robust, and interpretable detection of VSD from pediatric echocardiography videos. Its strong generalizability and seamless alignment with clinical workflows underscore its potential to enhance diagnostic capability in primary healthcare settings and to support early identification of congenital heart disease.
Liu et al. (Mon,) conducted a observational in Ventricular septal defect (VSD) (n=384). TSDNet model vs. State-of-the-art object detection models was evaluated on Video-level accuracy on the internal test set. The TSDNet model achieved 96.19% video-level accuracy and 86.86% frame-level accuracy for detecting ventricular septal defects in pediatric echocardiography videos on the internal test set.