In recent years, neural vocoders, an essential component of neural network-based speech synthesis systems, have achieved remarkable advances. However, a known issue in such systems is the degradation of synthesized speech quality due to an instantaneous amplitude decrement in voiced segments (hereinafter referred to as “dip"). In this study, we investigated methods for suppressing dips in synthesized waveforms generated by HiFi-GAN, a widely used representative neural vocoder. Our findings indicate that simply removing a single frame corresponding to the dip occurrence point, immediately after the first upsampling operation in HiFi-GAN, effectively suppresses the dip. Although this frame removal slightly shortens the synthesized waveform, the reduction in duration is 1.45 ms under typical settings, which is below the perceptual threshold for detecting a change in length. Furthermore, we found that restricting the synthesis range to approximately 200 ms centered around the dip is also sufficient to suppress the dip. This suggests that the removed frame itself may not be the direct cause of the dip, but rather that dips may be induced by extending the synthesis range.
Hirai et al. (Wed,) studied this question.