• Spanish and Italian available emotion recognition datasets and their performance. • In-the-wild emotion recognition in different languages and datasets. • Comparison of emotion recognition techniques as DeepSpectrum and wav2vec 2.0 and their performance. Within the context of developing new Socially Assistive Robots, emotion recognition has become a key factor, as it allows the robot to adapt to the user’s emotional state in real-world conditions. In this work, we focused on the analysis of Spanish and Italian voice-recording datasets, which, despite being widely spoken languages, do not receive the same level of attention as English. We selected five datasets: ELRA-S0329, EmoMatchSpanishDB, Emozionalmente, DEMoS, and EMOVO. Specifically, we centered our work on paralanguage—i.e., the vocal characteristics that accompany the message and clarify its meaning. We proposed the use of the DeepSpectrum method, which involves extracting a visual representation of the audio tracks and feeding them into a pretrained CNN model, as well as wav2vec 2.0, a self-supervised transformer module. We produced a total of four models and trained all of them on Spanish and Italian, finding that our DS-AM model (a VGG16-based model with two attention modules) and the wav2vec 2.0-based model outperformed the state-of-the-art (SOTA) models across all datasets. Finally, we performed an in-the-wild experiment, training the models on one dataset and testing them on another to simulate real-world conditions and assess how biased the models were toward specific datasets. We also included other datasets in the testing phase—CAVES (Cantonese), SUBESCO (Bangla), and RAVDESS (English)—to evaluate how the models perform in entirely different languages.
Ortega-Beltrán et al. (Sun,) studied this question.