Early detection of neurocognitive and mental health alterations is limited by the high costs and invasiveness of the protocols. Vocal biomarkers offer a non-invasive alternative, capturing prosodic and spectral features linked to cognitive and emotional states. Singing non-lexical syllabic vocalizations as a source for vocal biomarkers offers advantages over traditional speech-based methods, including language independence, scalability, and privacy preservation. This study evaluates whether singing melodies based on non-lexical syllables could provide reliable biomarkers sensitive to stress and cognitive load. Fourteen native Spanish-speaking participants completed a singing task under different stress conditions (acute stress vs. neutral) and different cognitive load conditions (immediate vs. delayed reproduction). Prosodic and spectral features were extracted, and results were explored in a series of linear mixed-effects models. Acute stress increased F0, decreased F1 and F2, and redistributed spectral energy (higher centroid, spread, flatness), producing a noisier output. Increased cognitive load led to shorter singing durations, increased F0, jitter and shimmer, and a flatter spectrum (lower kurtosis, higher flatness). Results demonstrated that stress engages autonomic arousal and articulatory changes, whereas cognitive load affects mainly control and stability. Shared markers index general arousal, while stress- and delay-specific features provide complementary sensitivity. Overall, singing non-lexical syllables stands as a feasible, low-burden, language-independent task for scalable vocal biomarker research. Alongside research on deep learning models, this opens new avenues for creating computer-based systems, either semi or fully automated, that can efficiently detect cognitive and emotional states.
Rodrigo-Herrero et al. (Fri,) studied this question.