PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 2, 2026Applied Sciences0 citationsOpen Access

Dimensional Emotion-Guided Conditional Modulation for Context-Aware Multimodal Driver Affect Recognition

View Full Paper
WSWei ShenXMXingang MouJYJing Yi

Key Points

  • This study aims to develop a framework that improves driver emotion recognition by integrating vehicle and visual data effectively.
  • Introduced a Dimensional Emotion-Guided Multi-task (DEGM) framework for emotion recognition.
  • Mapped vehicle data into a Valence–Arousal–Dominance (VAD) space to characterize emotions.
  • Utilized multi-task learning to optimize emotion classification and regression simultaneously.
  • Achieved an accuracy of 87.50% and a weighted F1-score of 0.8727 on the PPB driving emotion dataset.
  • Demonstrated effective cross-modal interaction leading to robust emotion detection.
  • Validated the framework's potential for real-world applications in intelligent transportation systems.

Abstract

Driver emotion recognition constitutes a fundamental pillar of intelligent cockpit systems, playing a pivotal role in enhancing driving safety and optimizing human–machine interaction. Despite the integration of vehicle sensor data in recent multimodal approaches, conventional fusion paradigms frequently encounter performance degradation due to the inherent noise and weak semantic correlation between vehicle telemetry and emotional states. To address these challenges, this study introduces a Dimensional Emotion-Guided Multi-task (DEGM) framework, a novel architecture designed to explicitly formalize the asymmetric roles of visual and vehicular modalities. Rather than employing simplistic feature concatenation, the proposed method maps multivariate vehicle data into a continuous Valence–Arousal–Dominance (VAD) space to characterize latent emotional tendencies within specific driving contexts. These predicted dimensions subsequently serve as semantic priors to conditionally modulate global facial representations through a Feature-wise Linear Modulation (FiLM) mechanism, facilitating robust and interpretable cross-modal interaction. Furthermore, the framework adopts a multi-task learning strategy that jointly optimizes discrete emotion classification and continuous dimension regression, leveraging the latter as a structural regularizer to refine the latent feature space. Comprehensive evaluations on the public PPB driving emotion dataset demonstrate that the proposed DEGM achieves a competitive accuracy of 87.50% and a weighted F1-score of 0.8727. The results validate that our framework provides a lightweight and robust paradigm for context-aware affect sensing, demonstrating strong potential for practical deployment in intelligent transportation systems.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Shen et al. (2026) studied this question.

synapsesocial.com/papers/69f5943c71405d493affefd0https://doi.org/10.3390/app16094312
Ask AI
Helpful
Bookmark
Share
View Full Paper