PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 10, 2026Nippon Onkyo Gakkaishi/Acoustical science and technology/Nihon Onkyo Gakkaishi0 citationsOpen Access

Speech Synthesis from Real-time MRI Articulatory Data with Precise F0 Reproduction

YOYuto OtaniSSShun SawadaHOHidefumi Ohmura

Key Points

  • The aim is to develop a speech synthesis method that effectively reproduces fundamental frequency using real-time MRI data.
  • Utilized real-time magnetic resonance imaging (rtMRI) to capture articulatory movements.
  • Developed an EfficientNetV2-BiLSTM network for F0-related feature extraction.
  • Employed a HiFi-GAN vocoder for generating high-fidelity waveforms.
  • Evaluated speech synthesis performance using the ATR 503 sentences rtMRI database.
  • Demonstrated intelligible speech synthesis with accurate F0 reproduction.
  • Confirmed that F0 can be estimated from single rtMRI frames without needing a temporal context.
  • Identified upward/forward larynx and tongue shifts corresponding to increasing F0.

Abstract

This paper presents a novel approach for speech synthesis using articulatory movements captured by real-time magnetic resonance imaging (rtMRI), focusing on fundamental frequency (F0) estimation mechanisms. Although recent rtMRI-based methods have achieved promising results, it remains unclear how F0 information is reproduced, given rtMRI's limited ability to capture vocal fold vibrations. To address this gap, we propose a speech synthesis method that processes only four consecutive rtMRI frames (~150 ms)—preventing reliance on extended linguistic context to infer F0. Our method employs an EfficientNetV2-BiLSTM network that enables sophisticated F0-related feature extraction for mel-spectrogram estimation, followed by a HiFi-GAN vocoder for high-fidelity waveform generation. Evaluations on the ATR 503 sentences rtMRI database demonstrate intelligible speech synthesis with accurate F0 reproduction. Building on these results, we further estimate F0 from single MRI frames, confirming that F0 can be derived without temporal context. To explore the underlying basis, we apply optical flow analysis to visualize subtle articulatory differences associated with F0 control, primarily revealing upward/forward larynx and tongue shifts with increasing F0. Additionally, distinct patterns were observed in male speakers at low F0 ranges. These findings empirically validate the relationship between articulatory configurations and F0 control, demonstrating feasibility in rtMRI-based speech synthesis.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Otani et al. (2026) studied this question.

synapsesocial.com/papers/69d892d16c1944d70ce04147https://doi.org/10.1250/ast.e25.75
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Toward Non-Invasive Voice Restoration: A Deep Learning Approach Using Real-Time MRI2025
  2. 2Speech synthesis via vocal tract shape extracted from real-time MRI2025
  3. 3Interpretable Modeling of Articulatory Temporal Dynamics from Real-Time MRI for Phoneme Recognition2026
  4. 4State-of-the-art speech production MRI protocol for new 0.55 Tesla scanners2024 · 1 citations
  5. 5Direct Speech Synthesis from Non-Invasive, Neuromagnetic Signals2024 · 2 citations