PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
July 21, 2025Expert Systems8 citations

A Comprehensive Review of Unimodal and Multimodal Emotion Detection: Datasets, Approaches, and Limitations

View Full Paper
PTPriyanka ThakurNKNirmal KaurNANaveen Aggarwal

Key Points

  • MAIN FINDING: Integrating multimodal data significantly enhances emotion detection accuracy using deep learning methods.
  • KEY EVIDENCE: Models trained on synchronized audio and video datasets show improved performance in recognizing emotions.
  • APPROACH: The review covers traditional machine learning and state-of-the-art deep learning techniques in emotion detection.
  • SIGNIFICANCE: This resource guides researchers in selecting the best datasets and models for advanced emotional AI applications.

Abstract

ABSTRACT Emotion detection from face and speech is inherent for human–computer interaction, mental health assessment, social robotics, and emotional intelligence. Traditional machine learning methods typically depend on handcrafted features and are primarily centred on unimodal systems. However, the unique characteristics of facial expressions and the variability in speech features present challenges in capturing complex emotional states. Accordingly, deep learning models have been substantial in automatically extracting intrinsic emotional features with greater accuracy across multiple modalities. The proposed article presents a comprehensive review of recent progress in emotion detection, spanning from unimodal to multimodal systems, with a focus on facial and speech modalities. It examines state‐of‐the‐art machine learning, deep learning, and the latest transformer‐based approaches for emotion detection. The review aims to provide an in‐depth analysis of both unimodal and multimodal emotion detection techniques, highlighting their limitations, popular datasets, challenges, and the best‐performing models. Such analysis aids researchers in judicious selection of the most appropriate dataset and audio‐visual emotion detection models. Key findings suggest that integrating multimodal data significantly improves emotion recognition, particularly when utilising deep learning methods trained on synchronised audio and video datasets. By assessing recent advancements and current challenges, this article serves as a fundamental resource for researchers and practitioners in the field of emotional AI, thereby aiding in the creation of more intuitive and empathetic technologies.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Thakur et al. (2025) studied this question.

synapsesocial.com/papers/689a0614e6551bb0af8cd5edhttps://doi.org/10.1111/exsy.70103
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1A Systematic Review of Emotion Recognition: From Unimodal Signals to Multimodal Integration2026
  2. 2Artificial Intelligence Approaches For Multimodal Emotion Understanding2024
  3. 3A Review on Transformer-Based Deep Learning Models for Multimodal Emotion Recognition2024
  4. 4A conceptual framework for deep learning-based multimodal emotion detection using facial expressions and physiological signals2026
  5. 5Enhancing Emotion Recognition through Multimodal Systems and Advanced Deep Learning Techniques2024