PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 15, 2026IEICE Transactions on Information and Systems0 citationsOpen Access

End-to-End Spontaneous Speech Recognition Based on Disfluency Labeling

KHKoharu HoriiMFMeiko FukudaKOKengo OHTA

Key Points

  • The aim is to enhance automatic speech recognition by addressing errors caused by disfluent utterances.
  • Proposed disfluency labeling to categorize segments of speech as fillers or hesitations.
  • End-to-end training of the ASR model incorporating labeled disfluency data.
  • Evaluation experiments comparing the new method against previous ASR techniques.
  • The disfluency labeling method achieved significantly higher recognition accuracy.
  • Explicit learning of disfluency features as labels improved the model’s ability to interpret spontaneous speech.

Abstract

Disfluent utterances in spontaneous speech, such as fillers and hesitations, cause recognition errors during automatic speech recognition (ASR). To address this problem, we propose a method called “disfluency labeling”, which replaces disfluent segments in transcription data with one of two labels: # (filler) or @ (hesitation). End-to-end training of the ASR model with such labeled data enables recognition of these disfluent segments as targets, like characters, allowing the extraction of what the speaker intended to say. In evaluation experiments, the proposed disfluency labeling method achieved higher recognition accuracy than the previous proposed ASR methods treating disfluencies, suggesting that explicit learning of disfluency features as labels is effective for improving spontaneous speech recognition.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Horii et al. (2026) studied this question.

synapsesocial.com/papers/69b64c67b42794e3e660da8fhttps://doi.org/10.1587/transinf.2025edp7157
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1On Disfluency and Non-lexical Sound Labeling for End-to-end Automatic Speech Recognition2024 · 3 citations
  2. 2Rich speech signal: exploring and exploiting end-to-end automatic speech recognizers’ ability to model hesitation phenomena2024
  3. 3Inclusive ASR for Disfluent Speech: Cascaded Large-Scale Self-Supervised Learning with Targeted Fine-Tuning and Data Augmentation2024 · 13 citations
  4. 4Inclusive ASR for Disfluent Speech: Cascaded Large-Scale Self-Supervised Learning with Targeted Fine-Tuning and Data Augmentation2024
  5. 5Artificial disfluency detection, uh no, disfluency generation for the masses2024 · 1 citations