PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 13, 20241 citationsOpen Access

Transcription-Free Fine-Tuning of Speech Separation Models for Noisy and Reverberant Multi-Speaker Automatic Speech Recognition

View Full Paper
WRWilliam RavenscroftGCGeorge CloseSGStefan Goetze

Key Points

Key points are not available for this paper at this time.

Abstract

One solution to automatic speech recognition (ASR) of overlapping speakers is to separate speech and then perform ASR on the separated signals. Commonly, the separator produces artefacts which often degrade ASR performance. Addressing this issue typically requires reference transcriptions to jointly train the separation and ASR networks. This is often not viable for training on real-world in-domain audio where reference transcript information is not always available. This paper proposes a transcription-free method for joint training using only audio signals. The proposed method uses embedding differences of pre-trained ASR encoders as a loss with a proposed modification to permutation invariant training (PIT) called guided PIT (GPIT). The method achieves a 6.4% improvement in word error rate (WER) measures over a signal-level loss and also shows enhancement improvements in perceptual measures such as short-time objective intelligibility (STOI).

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ravenscroft et al. (2024) studied this question.

synapsesocial.com/papers/68e64f88b6db6435875e005ehttps://doi.org/10.48550/arxiv.2406.08914
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Transcription-Free Fine-Tuning of Speech Separation Models for Noisy and Reverberant Multi-Speaker Automatic Speech Recognition2024 · 3 citations
  2. 2PixIT: Joint Training of Speaker Diarization and Speech Separation from Real-world Multi-speaker Recordings2024 · 1 citations
  3. 3PixIT: Joint Training of Speaker Diarization and Speech Separation from Real-world Multi-speaker Recordings2024 · 15 citations
  4. 4Improving Generalization of Speech Separation in Real-World Scenarios: Strategies in Simulation, Optimization, and Evaluation2024
  5. 5Improving Generalization of Speech Separation in Real-World Scenarios: Strategies in Simulation, Optimization, and Evaluation2024 · 3 citations