Previously we developed a voice production inversion system that predicts how speakers modulate their vocal fold physiology and subglottal pressure from the produced voice, toward ambulatory monitoring of vocal health and early detection of unhealthy vocal behavior. While this neural network was shown to predict changes in subglottal pressure and vocal fold geometry with reasonable accuracy, the neural network provides only point estimate predictions, but no information on the uncertainty of the predictions. Uncertainty information is essential to the interpretation of the predictions and clinical decision-making process. The goal of this study is to address this limitation. Two neural networks, a Bayesian neural network and a deep ensemble of neural networks, are developed to predict changes in vocal physiology as well as the confidence intervals of the predictions. The performance of the two neural networks is evaluated against running speech data from humans. Both methods are able to predict the subglottal pressure with reasonable accuracy and qualitatively predict the alternating adduction/abdcution during vowel-consonant transitions. The deep ensembles appear to slightly outperform the Bayesian neural network, in terms of smaller mean absolute errors and the narrower confidence intervals. Work supported by NIH.
Zhaoyan Zhang (2025) studied this question.