This study examines the progression from auditory-based speech intelligibility prediction to models incorporating spectrotemporal modulation features, and ultimately to machine learning models, from an engineering perspective. Prediction models based on auditory characteristics are well-suited for evaluating speech enhancement technologies in hearing aids and other auditory assistive devices. Traditional models based on auditory filterbank and modulation filterbank have provided a foundation for understanding speech perception, but the introduction of spectrotemporal modulation may enhance predictive accuracy. The presentation first introduces speech intelligibility prediction models that utilize frequency and temporal information analysis mechanisms, such as auditory and modulation filterbanks. Next, prediction models that employ spectrotemporal modulation and machine learning-based models using these as features are discussed. Finally, the presentation highlights trends in the Clarity Project, which leverages large-scale listening experiment data for hearing-impaired listeners, along with recent state-of-the-art approaches based on speech foundation models. These studies are expected to contribute not only to the development of engineering theoretical models but also to practical applications in hearing assistive technology and auditory diagnostics.
Katsuhiko Yamamoto (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: