Multimodal emotion identification based on audio-video data has attracted significant attention due to its outstanding results. Automated emotion recognition has grown in popularity during the last two decades. Because of the rapid advancement of Artificial Intelligence, discerning emotions in video has become a serious topic in human-computer interaction. The majority of known emotion identification systems only employ a single channel. However, in real life, people often disguise their true feelings, leading to relatively poor single-modal emotion detection accuracy. The Audio-Visual Residual Self-Attention Network (AV-RSANet) introduces a novel approach to multimodal emotion recognition by integrating both audio and visual inputs, effectively compensating for the limitations of a single modality and enhancing the accuracy of emotion detection in video content. More specifically, we offer cutting-edge feature-extraction networks tailored for video and audio. A multimodal emotion identification model is then produced by combining audio and visual inputs at the model level. The benchmark multimodal dataset, Ryerson's Audio-Visual Database of Emotional Speech and Song (RAVDESS), is utilised to assess the performance of the suggested models. The proposed models on the RAVDESS dataset achieve high projected accuracies of 99.8%. Furthermore, head-to-head comparisons with cutting-edge emotion identification algorithms corroborate the models' performance. • Proposes AV-RSANet for efficient audio–visual emotion recognition. • Introduces residual self-attention for cross-modal feature fusion. • Preserves unimodal saliency while enabling dynamic cross-modal interaction. • Achieves 99.8% accuracy on the RAVDESS dataset. • Demonstrates real-time performance with a lightweight architecture.
Jothimani et al. (Fri,) studied this question.