PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 18, 202415 citationsOpen Access

PixIT: Joint Training of Speaker Diarization and Speech Separation from Real-world Multi-speaker Recordings

View Full Paper
JKJoonas KaldaCPClément PagèsRMRicard Marxer

Key Points

Key points are not available for this paper at this time.

Abstract

A major drawback of supervised speech separation (SSep) systems is their reliance on synthetic data, leading to poor realworld generalization.Mixture invariant training (MixIT) was proposed as an unsupervised alternative that uses real recordings, yet struggles with over-separation and adapting to longform audio.We introduce PixIT, a joint approach that combines permutation invariant training (PIT) for speaker diarization (SD) and MixIT for SSep.With a small extra requirement of needing SD labels during training, it solves the problem of over-separation and allows stitching local separated sources leveraging existing work on clustering-based neural SD.We measure the quality of the separated sources via applying automatic speech recognition (ASR) systems to them.PixIT boosts the performance of various ASR systems across two meeting corpora both in terms of the speaker-attributed and utterancebased word error rates while not requiring any fine-tuning.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Kalda et al. (2024) studied this question.

synapsesocial.com/papers/68e643d5b6db6435875d537ahttps://doi.org/10.21437/odyssey.2024-17
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1EEND-M2F: Masked-attention mask transformers for speaker diarization2024 · 18 citations
  2. 2Powerset multi-class cross entropy loss for neural speaker diarization2023 · 115 citations
  3. 3End-to-End Neural Speaker Diarization with Permutation-Free Objectives2019 · 241 citations
  4. 4GPU-accelerated Guided Source Separation for Meeting Transcription2023 · 29 citations
  5. 5The AMI Meeting Corpus: A Pre-announcement2006 · 926 citations