PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 22, 20240 citationsOpen Access

Multimodal Segmentation for Vocal Tract Modeling

View Full Paper
RJRishi JainBYBohan YuPWPeter Wu

Key Points

Key points are not available for this paper at this time.

Abstract

Accurate modeling of the vocal tract is necessary to construct articulatory representations for interpretable speech processing and linguistics. However, vocal tract modeling is challenging because many internal articulators are occluded from external motion capture technologies. Real-time magnetic resonance imaging (RT-MRI) allows measuring precise movements of internal articulators during speech, but annotated datasets of MRI are limited in size due to time-consuming and computationally expensive labeling methods. We first present a deep labeling strategy for the RT-MRI video using a vision-only segmentation approach. We then introduce a multimodal algorithm using audio to improve segmentation of vocal articulators. Together, we set a new benchmark for vocal tract modeling in MRI video segmentation and use this to release labels for a 75-speaker RT-MRI dataset, increasing the amount of labeled public RT-MRI data of the vocal tract by over a factor of 9. The code and dataset labels can be found at rishiraij. github. io/multimodal-mri-avatar/.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Jain et al. (2024) studied this question.

synapsesocial.com/papers/68e63c0bb6db6435875cd985https://doi.org/10.48550/arxiv.2406.15754
Ask AI
Helpful
Bookmark
Share
View Full Paper