PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 8, 20250 citationsOpen Access

Length Aware Speech Translation for Video Dubbing

View Full Paper
HCHarveen Singh ChadhaASAswin Shanmugam SubramanianVJVikas Joshi

Key Points

  • Length-sensitive speech translation significantly improves synchronization quality in video dubbing.
  • The proposed model achieved a mean opinion score gain of 0.65 for Korean and 0.34 for Spanish translations.
  • Length-aware beam search enables various translation lengths to be produced in one decoding pass.
  • Maintaining BLEU scores comparable to a baseline indicates effective translation without compromising quality.

Abstract

In video dubbing, aligning translated audio with the source audio is a significant challenge. Our focus is on achieving this efficiently, tailored for real-time, on-device video dubbing scenarios. We developed a phoneme-based end-to-end length-sensitive speech translation (LSST) model, which generates translations of varying lengths short, normal, and long using predefined tags. Additionally, we introduced length-aware beam search (LABS), an efficient approach to generate translations of different lengths in a single decoding pass. This approach maintained comparable BLEU scores compared to a baseline without length awareness while significantly enhancing synchronization quality between source and target audio, achieving a mean opinion score (MOS) gain of 0.34 for Spanish and 0.65 for Korean, respectively.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Chadha et al. (2025) studied this question.

synapsesocial.com/papers/68e6f342f8145af55aeacad0https://doi.org/10.48550/arxiv.2506.00740
Ask AI
Helpful
Bookmark
Share
View Full Paper