PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 16, 2026Sensors0 citationsOpen Access

Facial Expression Annotation and Analytics for Dysarthria Severity Classification

SDShufei DuanYGYuxin GuoLFLonghao Fu

Key Points

  • This research aims to develop a multimodal framework for classifying the severity of dysarthria by integrating facial and acoustic features.
  • Designed a multi-level annotation algorithm to address data scarcity.
  • Modeled facial topology using Delaunay triangulation and graph convolutional networks.
  • Quantified abnormal muscle coordination with facial action units.
  • Proposed a framework for fusing multimodal features for better disease classification.
  • Achieved an accuracy rate of 92.0% using the THE-POSSD dataset.
  • F1 score of 91.6%, significantly exceeding baseline single-modality results.
  • Revealed insights into facial movements affected by emotional states.
  • Verified the compensatory role of visual patterns in assessing auditory challenges.

Abstract

Dysarthria in patients post-stroke is often accompanied by central facial paralysis, which impairs facial motor control and emotional expression. Current assessments rely on acoustic modalities, overlooking facial pathological cues and their correlation with emotional expression, which hinders comprehensive disease assessment. To address this issue, we propose a multimodal severity classification framework that integrates facial and acoustic features. Firstly, a multi-level annotation algorithm based on a pre-trained model and motion amplitude was designed to overcome the problem of data scarcity. Secondly, facial topology was modeled using Delaunay triangulation, with spatial relationships captured via graph convolutional networks (GCNs), while abnormal muscle coordination is quantified using facial action units (AUs). Finally, we proposed a multimodal feature set fusion technology framework to achieve the compensation of facial visual features for acoustic modalities and the analysis of disease classification. Our experimental results using the THE-POSSD dataset demonstrate an accuracy of 92.0% and an F1 score of 91.6%, significantly outperforming single-modality baselines. This study reveals the changes in facial movements and sensitive areas of patients under different emotional states, verifies the compensatory ability of visual patterns for auditory patterns, and demonstrates the potential of this multimodal framework for objective assessment and future clinical applications in speech disorders.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Duan et al. (2026) studied this question.

synapsesocial.com/papers/6992b45f9b75e639e9b09443https://doi.org/10.3390/s26041239
Ask AI
Helpful
Bookmark
Share
View Full Paper