Acoustic features have been widely used to differentiate between individuals with and without depression. While spectral features describe the energy distribution in the speech signal, prosodic features capture suprasegmental patterns such as rhythm, intonation, and stress, which are key indicators of emotional and cognitive states. This study analyzed the acoustic features of speech produced by people of varying levels of depression severity,aiming to identify the relation between speech production and depression severity. Using Praat, I extracted prosodic features including time-normalized fundamental frequency, pitch range, speaking rate, and percentage pause time as well as spectral features such as power spectral density and mel-frequency cepstral coefficients. The results showed that greater percent pause time was correlated with elevated anxiety levels, while reduced F0 range and pitch variability were associated with increased depression severity. Additionally, slower speech rate exhibited a strong negative correlation with depression severity. MFCC values also exhibited negative correlations with depression severity, reflecting reduced vocal dynamics in depressed speech. These findings suggest that acoustic features may provide objective cues for the evaluation of depression severity, providing a supplementary tool for mental health assessment.
Shuqi Huang (Wed,) studied this question.