Speech that stands out from background noise is called a “pop-out” voice. Several acoustic features are known to contribute to the degree of pop-out, including fundamental frequency, power in the frequency band above 1 kHz, and dynamic feature (DF). Our research focuses on controlling the degree of pop-out in Text-to-Speech (TTS) by modifying relevant acoustic features, using JETS, which is an end-to-end TTS model, as the base. We first attempted to incorporate DF into the model by adding a “DF predictor” alongside the existing pitch, energy, and duration predictors in the variance adaptor of JETS. Preliminary evaluations of the model without explicit DF control showed that adding the DF predictor did not degrade the subjective quality of the synthetic samples, and the mean DF values of the synthetic samples were comparable to those of the original recordings. To produce more pop-out-like samples, we introduced an approach that emphasizes the output of the DF predictor. This method generated samples with higher DF when the emphasis weighting was set to a large value. However, some degradation was observed in the generated speech. We are currently working to address this issue.
Uesugi et al. (2025) studied this question.