PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 22, 2026Computers, materials & continua/Computers, materials & continua (Print)0 citationsOpen Access

SYMPHONIA–Enhanced Multimodal Emotion Recognition with Dual-Branch Dynamic Attention and Hierarchical Adaptive Fusion

AAAkmalbek AbdusalomovMMMukhriddin MukhiddinovKAKamola Abdurashidova

Key Points

  • The aim is to improve emotion recognition by integrating visual and textual modalities using innovative architecture.
  • Developed SYMPHONIA architecture with Facial Emotion and Textual Emotion branches.
  • Utilized Vision Transformers for facial expressions and RoBERTa embeddings for language.
  • Implemented a Dual-Branch Dynamic Attention Mechanism and Hierarchical Adaptive Fusion Module.
  • Achieved 80.9% accuracy and 80.1% F1-score on IEMOCAP dataset, outperforming competitors.
  • Secured 74.2% accuracy and 73.5% F1-score on MELD dataset.
  • Demonstrated generalization with 66.9% accuracy across datasets.

Abstract

Human emotions are intricate and difficult to decipher through various modalities. Current methodologies frequently employ inflexible fusion strategies that do not consider the dynamic and context-sensitive characteristics of emotional expressions in both visual and textual mediums. This paper presents SYMPHONIA (Synchronizing Facial and Textual Modalities for Emotion Understanding), an innovative architecture engineered to capture and amalgamate emotional signals from facial expressions and language, attuned to contextual and modality interactions. There are two parts to SYMPHONIA: a Facial Emotion Branch that uses Vision Transformers and facial landmarks, and a Textual Emotion Branch that uses RoBERTa embeddings and graph-based reasoning. A Dual-Branch Dynamic Attention Mechanism and a Hierarchical Adaptive Fusion Module are used to connect these branches. SYMPHONIA beat the best models on four datasets: IEMOCAP, MELD, CMU-MOSI, and CMU-MOSEI. It got 80.9% accuracy and 80.1% F1-score on IEMOCAP, which was better than Dualgats (74.8%) and EmoCLIP (75.3%). SYMPHONIA got 74.2% accuracy and 73.5% F1-score for MELD. It beat its competitors by getting a 0.86 Pearson correlation on MOSI and a 0.83 on MOSEI for predicting sentiment. Cross-dataset tests showed that SYMPHONIA could generalize, with 66.9% accuracy when trained on IEMOCAP and tested on MELD. This was better than all the baselines. These results show that SYMPHONIA is good at recognizing emotions and analyzing sentiment in different situations, which shows that it can adapt and do well in different settings.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Abdusalomov et al. (2026) studied this question.

synapsesocial.com/papers/69e866f16e0dea528ddeb4e1https://doi.org/10.32604/cmc.2026.077057
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1A Complementary Fusion Framework for Robust Multimodal Emotion Recognition2025
  2. 2Multimodal emotion recognition based on a fusion of audiovisual information with temporal dynamics2024 · 38 citations
  3. 3A deep learning framework for emotion recognition in music using multimodal data fusion2026
  4. 4Cross-modal Synergy for Enhancing Emotion Recognition Through Integrated Audio–Video Fusion Techniques2025
  5. 5ECMF: Enhanced Cross-Modal Fusion for Multimodal Emotion Recognition in MER-SEMI Challenge2025