Parallel end-to-end TTS models have made notable progress in speech naturalness but still suffer from limitations in fluency and computational efficiency. This letter presents GSViTS, a parallel TTS framework that achieves robust alignment and enchanced prosody by integrating a differentiable SoftDTW-based alignment module with uncertainty-aware duration prediction. In addition, a Dual-Path Gated Linear Attention (DP-GLA) mechanism is introduced to support efficient long-sequence modeling with reduced computational overhead. Experiments on LJSpeech and VCTK demonstrate 20% MOS improvement and 7% lower computational cost over baselines.
Zhao et al. (Thu,) studied this question.