PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 14, 2026The Journal of the Acoustical Society of America0 citations

Speaker segmentation and diarization using graph transformers on the fearless steps Apollo 11 corpus

View Full Paper
MSMeena Chandra ShekarJHJohn H. Hansen

Key Points

  • This research aims to improve speaker diarization by addressing limitations in traditional datasets and utilizing advanced modeling techniques.
  • Leveraged LLMs for initial speaker segmentation labels based on conversational structure and lexical cues.
  • Employed a semi-supervised graph transformer architecture to enhance speaker embeddings.
  • Final clustering performed using agglomerative hierarchical clustering with 80 h of labeled audio from five channels and 282 unique speakers.
  • Achieved improved speaker segmentation and diarization in complex acoustic environments.
  • Demonstrated effective handling of overlapping speech and speaker variability.
  • Validation shows significant enhancement in clustering accuracy over traditional methods.

Abstract

Speaker diarization is the task of partitioning an audio stream into segments with the objective of determining “who spoke when.” Traditionally, speaker diarization systems have relied on datasets that typically feature a limited number of speakers, ranging from 2–6 per session often consist of clean or simulated speech. These datasets do not reflect the speaker variability, overlapping speech, or channel diversity observed in real-world operations. In our study, we use the Fearless Steps Apollo-11 (FS-A11) corpus, which presents a wide range of speakers, diverse acoustic environments, varying speaker utterance durations, and an unknown total number of speakers. To address these challenges, we leverage LLMs to generate initial speaker segmentation labels. LLMs are capable of analyzing conversational structure, initiation patterns, and lexical cues specific to Apollo communications. These labels are refined using variable word-based segmentation. Finally, we propose a semi-supervised graph transformer architecture for speaker diarization to enhance speaker embeddings by modeling relationships between segments, taking into account both temporal and contextual information. Final clustering is performed using agglomerative hierarchical clustering (AHC). The training set includes 80 h of labeled audio from five channels, with 282 unique speakers. Evaluation and test sets each consist of 10 h.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Shekar et al. (2025) studied this question.

synapsesocial.com/papers/6a0567fda550a87e60a2055chttps://doi.org/10.1121/10.0041563
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Apollo’s Unheard Voices: Graph Attention Networks for Speaker Diarization and Clustering for Fearless Steps Apollo Collection2024 · 3 citations
  2. 2Online Speaker Diarization of Meetings Guided by Speech Separation2024 · 8 citations
  3. 3Robust Target Speaker Diarization and Separation via Augmented Speaker Embedding Sampling2025
  4. 4End-to-End Sequence Labeling for Myanmar Speaker Diarization2024
  5. 5A Hybrid DMFCC-LPC Based Feature Extraction with DCNN Clustering for Speaker Diarization2024 · 1 citations