Multimodal large language models demonstrated limited diagnostic utility for detecting clinically significant pediatric ECG abnormalities, with ChatGPT achieving the highest positive likelihood ratio of 2.05.
Cross-Sectional (n=264)
Single-blind
No
Does multimodal LLM interpretation improve diagnostic accuracy for clinically significant and emergency abnormalities in pediatric patients undergoing ECG evaluation?
Current multimodal LLMs show limited diagnostic utility in pediatric ECG interpretation and cannot be used as standalone diagnostic tools, though they may serve as adjunctive screening aids under clinician oversight.
Effect estimate: +LR 2.05 (95% CI 1.42-2.96)
• This is theFCA1 first head-to-head comparative diagnostic accuracy study of multimodal LLMs in pediatric ECG evaluation, using likelihood ratios as primary outcome measures. • All three LLMs showed limited rule-in utility (+LR near 1.0); Gemini achieved potentially meaningful rule-out performance for emergency arrhythmias (-LR = 0.07), but with wide confidence intervals reflecting the small emergency subgroup (n = 22). • Gemini's 100% sensitivity in the emergency subgroup reflects overcalling (specificity 30.2%) consistent with a triage/screening behavior rather than diagnostic precision.
Saraç et al. (Tue,) conducted a cross-sectional in Pediatric ECG abnormalities (n=264). Multimodal large language models (ChatGPT, Gemini 3, Microsoft Copilot) vs. Expert consensus (three pediatric cardiologists) was evaluated on Positive likelihood ratio (+LR) for clinically significant ECG abnormalities (+LR 2.05, 95% CI 1.42-2.96). Multimodal large language models demonstrated limited diagnostic utility for detecting clinically significant pediatric ECG abnormalities, with ChatGPT achieving the highest positive likelihood ratio of 2.05.