ABSTRACT Objective Laryngomalacia diagnosis and subtype classification is typically based on flexible nasolaryngoscopy, but this can be challenging, particularly for clinicians who see it infrequently. Infant movement may also prolong the procedure. Automated computer vision (CV) assessment of video nasolaryngoscopy may improve diagnostic accuracy, shorten procedure time, and reduce infant discomfort. The aim was to develop a CV model to segment laryngeal anatomy, differentiate between normal and laryngomalacia cases, and classify laryngomalacia type (1, 2, or 3). Methods A total of 241 pediatric nasolaryngoscopy videos were obtained from The Hospital for Sick Children between August 2023 and March 2025 from participants < 1 year old. Videos were annotated and randomly assigned to a training set 192 (80%) and testing set 49 (20%). Three binary classification models (fine‐tuned for feature extraction) were developed to predict laryngomalacia and classify its type. CV predictions were compared with diagnoses made by three staff otolaryngologists. Results The CV model's Dice similarity coefficient for laryngeal anatomy segmentation was 0.86 (95% CI: 0.85–0.88). The model differentiated normal from laryngomalacia cases with an accuracy of 0.90 and predicted type 1 and type 2 laryngomalacia with accuracies of 0.86 and 0.90, respectively. A type 3 model was not developed due to insufficient cases. Conclusion The CV model successfully segmented laryngeal anatomy, differentiated normal from laryngomalacia cases, and classified type 1 and 2 laryngomalacia from pediatric video nasolaryngoscopy. With dataset expansion, further algorithm refinement, and external validation, this model may serve as a clinical decision support tool in the future. Level of Evidence 3.
Madan et al. (Mon,) studied this question.