Machine learning models accurately assessed mitral regurgitation severity with a pooled AUROC of 0.97, sensitivity of 0.93, and specificity of 0.96 across nine studies.
Do machine learning algorithms improve diagnostic accuracy for assessing mitral regurgitation severity in adults compared to conventional echocardiographic interpretation?
Machine learning models demonstrate high diagnostic accuracy for assessing mitral regurgitation severity, though prospective standardized trials are needed for clinical adoption.
Accurate assessment of mitral regurgitation (MR) severity is crucial for guiding clinical management, but is often limited by the subjectivity and variability of traditional echocardiographic evaluations. Machine learning (ML) models offer potential for automated, objective MR grading, yet their diagnostic performance remains underexplored. This systematic review and meta-analysis aim to evaluate the diagnostic accuracy of ML-based models for assessing MR severity. We searched five different databases for studies evaluating ML algorithms (deep learning or traditional ML) for MR severity assessment in adults. Data were extracted and the risk of bias was assessed using the PROBAST+AI tool. A bivariate random-effects model was used to pool diagnostic metrics, with heterogeneity quantified via I 2 statistics and explored through meta-regression and subgroup analyses. Publication bias was evaluated using Deeks’ test and funnel plot. Nine studies met inclusion criteria, demonstrating strong ML performance with a pooled AUROC of 0.97 (95% CI: 0.96–0.98), sensitivity of 0.93 (95% CI: 0.83–0.97), and specificity of 0.96 (95% CI: 0.92–0.98). High heterogeneity (I 2 > 70%) was observed, partly explained by variations in validation methods and sample size. No significant publication bias was detected (Deeks’ p=0.64). The certainty of the evidence was moderate due to heterogeneity and the retrospective study design. ML models demonstrate good diagnostic accuracy for assessing MR severity, with the potential to enhance clinical decision-making by reducing subjectivity. However, high heterogeneity and limited external validation necessitate prospective, standardized trials to ensure generalizability and clinical adoption. • Machine learning models demonstrate high diagnostic accuracy for assessing mitral regurgitation severity, with pooled AUROC of 0.97 and sensitivity and specificity of 0.96. • Across nine included studies, ML approaches consistently showed strong performance compared with conventional echocardiographic interpretation. • Substantial heterogeneity was observed, partly driven by differences in validation strategies and dataset size. • Prospective studies with standardized methodologies and external validation are needed before routine clinical implementation.
Eini et al. (2026) studied this question. Machine learning models accurately assessed mitral regurgitation severity with a pooled AUROC of 0.97, sensitivity of 0.93, and specificity of 0.96 across nine studies.