PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 13, 2026SHILAP Revista de lepidopterología0 citationsOpen Access

Improving zero-shot style transfer text-to-speech by disentangled fine-grained style modeling

View Full Paper
EEEray ErenQLQingju LiuAAA. Alwan

Key Points

  • The aim is to enhance zero-shot style transfer in text-to-speech systems using smaller models without compromising quality.
  • Proposed a zero-shot method leveraging the GenerSpeech backbone and fine-grained style encoders.
  • Implemented a mutual-information minimization loss to separate speaker identities and styles.
  • Applied a maximum-mean-discrepancy-guided cycle consistency loss for better style embedding diversity.
  • Achieved a relative average style preference improvement of 31% over baseline methods.
  • Obtained a prosody similarity mean opinion score of 3.64 on the VCTK dataset.

Abstract

Recent zero-shot style-transfer speech synthesis methods have shown promising results and addressed adaptation to unseen speaking styles. While most state-of-the-art methods generalize to new speakers and styles using large models or corpora, achieving similar generalization with a smaller model remains an open challenge. We propose a zero-shot method that uses the small GenerSpeech backbone plus a fine-grained style encoder. To disentangle speakers, global/fine-grained styles, and content embeddings, we introduce a mutual-information minimization loss. To further disentangle style from speaker and boost style embedding diversity, we introduce a maximum-mean-discrepancy-guided cycle consistency loss. Experimental results show the proposed method outperforms baseline zero-shot style-transfer methods (GenerSpeech, YourTTS, VALL-E-X) with a relative average style preference improvement of 31% and a 3.64 prosody prosody similarity mean opinion score on VCTK.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Eren et al. (2026) studied this question.

synapsesocial.com/papers/69b3acd302a1e69014ccecechttps://doi.org/10.1121/10.0042974
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1NaturalSpeech: End-to-End Text-to-Speech Synthesis With Human-Level Quality2024 · 163 citations
  2. 2Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale2023 · 7 citations
  3. 3Cycle consistent network for end-to-end style transfer TTS training2021 · 31 citations
  4. 4Cross-Speaker Emotion Transfer by Manipulating Speech Style Latents2023 · 4 citations
  5. 5StyleTTS: A Style-Based Generative Model for Natural and Diverse Text-to-Speech Synthesis2025 · 33 citations