PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 14, 2026The Journal of the Acoustical Society of America0 citations

Estimating subglottal pressure and vocal fold adduction with confidence intervals using deep ensembles and Bayesian neural networks

View Full Paper
ZZZhaoyan Zhang

Key Points

  • This study aims to enhance the predictions of a voice production inversion system by incorporating uncertainty information.
  • Developed a Bayesian neural network and a deep ensemble of neural networks.
  • Evaluated performance using running speech data from humans.
  • Compared prediction accuracy and confidence intervals of both neural network methods.
  • Both neural networks accurately predicted subglottal pressure.
  • Deep ensembles produced smaller mean absolute errors and narrower confidence intervals compared to the Bayesian neural network.
  • Both methods qualitatively predicted vocal fold adduction/abduction during vowel-consonant transitions.

Abstract

Previously we developed a voice production inversion system that predicts how speakers modulate their vocal fold physiology and subglottal pressure from the produced voice, toward ambulatory monitoring of vocal health and early detection of unhealthy vocal behavior. While this neural network was shown to predict changes in subglottal pressure and vocal fold geometry with reasonable accuracy, the neural network provides only point estimate predictions, but no information on the uncertainty of the predictions. Uncertainty information is essential to the interpretation of the predictions and clinical decision-making process. The goal of this study is to address this limitation. Two neural networks, a Bayesian neural network and a deep ensemble of neural networks, are developed to predict changes in vocal physiology as well as the confidence intervals of the predictions. The performance of the two neural networks is evaluated against running speech data from humans. Both methods are able to predict the subglottal pressure with reasonable accuracy and qualitatively predict the alternating adduction/abdcution during vowel-consonant transitions. The deep ensembles appear to slightly outperform the Bayesian neural network, in terms of smaller mean absolute errors and the narrower confidence intervals. Work supported by NIH.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhaoyan Zhang (2025) studied this question.

synapsesocial.com/papers/6a0567bca550a87e60a1ff46https://doi.org/10.1121/10.0040196
Ask AI
Helpful
Bookmark
Share
View Full Paper