Within the past several years, technology has emerged to allow creation of highly natural-sounding synthetic “voices” that accurately mimic an individual human’s speech using brief samples of that person’s natural speech. Based on this technology, creation of personal synthetic voices for use in speech generating devices (SGDs) by people with neurodegenerative conditions such as ALS is becoming commonplace with multiple services now available that assist clients to record speech samples before their speech is heavily affected by disease progression. Children who must use an SGD for communication may also benefit from a personal synthetic voice for their SGD. However, such children typically cannot produce the normally articulated speech samples needed to train the underlying transformer models for the latest generation of Text to Speech (TTS) systems. In this study, we use transfer learning from a model based on a large diverse set of speakers to produce synthetic voices for young children using small collections of single words and short utterances (with and without articulation errors) as the training material. We present preliminary results in terms of intelligibility, audio quality, naturalness, and resemblance of voice quality to that of the target talker using sentence and paragraph-level samples for evaluation by listeners.
Lilley et al. (Wed,) studied this question.