PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 11, 2026Journal of King Saud University - Computer and Information Sciences0 citationsOpen Access

Multimodal Stress-Aware Ensemble Learning (MSAEL): tri-modal behavioral-physiological fusion for real-time proxy-based stress recognition integrating typing, facial, and voice dynamics

MKMousa Mohammed Khubrani

Key Points

  • This work aims to improve stress recognition by integrating multiple behavioral and physiological modalities.
  • Developed Multimodal Stress-Aware Ensemble Learning (MSAEL) for stress-proxy classification.
  • Utilized datasets including FER-2013, RAVDESS, and CMU Keystroke Dynamics for model training.
  • Implemented a stacked ensemble classifier with adaptive attention-driven fusion.
  • Achieved 91.4% accuracy and 0.956 AUC-ROC, outperforming other existing methods.
  • Validated unique contributions of each modality in stress detection with statistical analyses.
  • Demonstrated the framework's potential for applications in occupational health and telemedicine.

Abstract

The progression of Real-Time, stress-aware classification using high-arousal affective proxies derived from facial, vocal, and behavioral cues in affective computing is hindered by the fragmentation of human stress expression across emotional, physiological, and cognitive–behavioral domains. Unimodal and bimodal techniques generally fail to capture this multidimensional complexity, leading to limited applicability and unstable predictions in unrestrained settings. In this work, stress is modeled using high-arousal negative emotional states as proxies, and the proposed framework focuses on stress-proxy classification rather than direct physiological stress detection. This study describes Multimodal Stress-Aware Ensemble Learning (MSAEL), a technically sophisticated tri-modal system that incorporates facial emotion patterns, vocal–physiological indicators, and typing-driven behavioral dynamics into a deep ensemble architecture to address this breach. Therefore, to forecast anxiety across Low, Moderate, and High categories, the system uses modality-specific feature extractors, adaptive attention-driven fusion, and a stacked ensemble classifier. In this study, a PRISMA-guided methodological review of 30 multimodal emotional computing research papers informed this framework's methodology. Experimental datasets included FER-2013 (facial expressions), RAVDESS (speech emotion), and CMU Keystroke Dynamics, composed with synthetically aligned tri-modal samples to approximate real-time use. MSAEL surpassed unimodal, bimodal, and advanced multimodal methods, including MDFN, AVEC Fusion, and EAEL, with 91.4% accuracy, 91.3% F1-score, and 0.956 AUC-ROC. Extirpation findings validate each modality's distinct contribution, while correlation and statistical inference studies demonstrate the tri-modal fusion strategy's consistency, reliability, and complementarity. MSAEL is a vigorous stress detection framework that can overcome noise sensitivity, inter-modal discrepancies, and perplexing emotional overlaps. Multimodal affective computing is improved by demonstrating that behavioral, physiological, and expressive signs are most effective for identifying stress. The study demonstrates that MSAEL might be used in occupational well-being nursing, digital medicines, telemedicine, and edge-intelligent human–machine interactions. Unified tri-modal dataset development, physiological wearable integration, edge optimization, and privacy-preserving deployment via federated learning are future goals. Instead of performing subject-level, synchronously recorded, multimodal measurements, the proposed framework seeks to analyze modality-level complementarity using both the model-based methodology and a methodology based on labeled, aligned fusion of independently collected public datasets. Therefore, as on-site synchronized stress sensing of individuals within companies cannot be done, results should not be used to justify its real application. Instead, empirical data on inference latency should be used, as these two studies will not align.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Mousa Mohammed Khubrani (2026) studied this question.

synapsesocial.com/papers/6a0171ce3a9f334c28271d61https://doi.org/10.1007/s44443-026-00765-9
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Machine and Deep Learning Models for Stress Detection Using Multimodal Physiological Data2025 · 47 citations
  2. 2Multimodal Daily-Life Emotional Recognition Using Heart Rate and Speech Data From Wearables2024 · 12 citations
  3. 3Combining Multimodal Features within a Fusion Network for Emotion Recognition in the Wild2015 · 54 citations
  4. 4Expert Systems in Behavioral and Mental Healthcare: Applications of AI in Decision‐Making and Consultancy2022 · 8 citations
  5. 5Non invasive human stress detection using key stroke dynamics and pattern variations2013 · 45 citations