PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 29, 20260 citationsOpen Access

Attention-Based CNN-BiGRU for Speech Clarity Estimation in ALS

View Full Paper
PKPayal KatiyarRRRishika RajTRThilagavathy R

Key Points

  • The aim is to develop an automatic system for estimating speech clarity in patients with Amyotrophic Lateral Sclerosis.
  • Hybrid deep learning architecture combining CNN for feature extraction and BiGRU for temporal modeling.
  • Utilization of a self-attention mechanism to focus on essential speech segments.
  • Implementation of a soft labeling strategy based on various acoustic features.
  • Model trained and evaluated on the TORGO dysarthric speech dataset.
  • Achieved a Pearson correlation of 0.778.
  • Demonstrated a Mean Absolute Error (MAE) of 0.083.
  • Attained an R² score of 0.567, indicating effective performance.
  • Integrated into a web application for real-world usability.

Abstract

This project presents an end-to-end system for automatic speech clarity estimation in patients with Amyotrophic Lateral Sclerosis (ALS). The proposed approach models speech clarity as a continuous regression problem, enabling more realistic and clinically meaningful monitoring of speech degradation compared to traditional classification-based methods.The system utilizes a hybrid deep learning architecture combining Convolutional Neural Networks (CNN) for spectral feature extraction, Bidirectional Gated Recurrent Units (BiGRU) for temporal modeling, and a self-attention mechanism to focus on informative speech segments. Additionally, a soft labeling strategy based on acoustic features such as spectral properties, harmonic-to-noise ratio (HNR), and voiced ratio is employed to better capture the continuous nature of speech clarity.The model is trained and evaluated on the TORGO dysarthric speech dataset, which contains real-world variability in speakers and recording conditions. Experimental results demonstrate strong performance, achieving a Pearson correlation of 0.778, Mean Absolute Error (MAE) of 0.083, and an R² score of 0.567, highlighting the effectiveness of the proposed method.To enable real-world usability, the model is integrated into a full-stack web application (Speech Clarity Analyser) that allows users to record speech, obtain clarity scores, and monitor changes over time. The application is built using React (frontend), FastAPI (backend), and Supabase (database and authentication), ensuring secure data handling and scalable deployment.This work contributes to the development of non-invasive, automated tools for clinical monitoring of speech impairment, with potential applications in early diagnosis, remote patient monitoring, and assistive healthcare technologies.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Katiyar et al. (2026) studied this question.

synapsesocial.com/papers/69c8c30dde0f0f753b39da98https://doi.org/10.5281/zenodo.19265198
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Speech Clarity Analysis for ALS Patients Using CNN-BiGRU and Speaker-Relative Regression2026
  2. 2AI-Driven Early Detection of Amyotrophic Lateral Sclerosis (ALS): A Machine Learning Approach Using Acoustic Biomarkers2025 · 2 citations
  3. 3Meta-learning approach for ALS detection using sustained vowel phonations2026
  4. 4A Hybrid Machine Learning and Blockchain Architecture for Enhanced ALS Detection2025 · 1 citations
  5. 5Enhancing dysarthria severity classification: efficient audio based deep learning models2025 · 2 citations