PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 12, 2026The Journal of the Acoustical Society of America0 citations

Best practices in evaluating forced-alignment accuracy: The case of Mandarin varieties

View Full Paper
SLSuyuan LiuMSMárton SóskuthySZSijia Zhang

Key Points

  • The research aims to assess the accuracy of forced alignment in different Mandarin varieties and establish evaluation standards.
  • Evaluated machine-generated alignments from Montreal Forced Aligner (MFA).
  • Compared alignments against two sets of human baselines.
  • Used a Bayesian hierarchical multivariate regression model for analysis.
  • Examined differences in alignment accuracy across various sequence types.
  • Found closer agreement between human aligners compared to humans and MFA.
  • Identified large accuracy differences across different alignment sequences.
  • Noticed effects of speech rate and speaker variability.
  • Observed no variation in accuracy across different Mandarin varieties.

Abstract

Forced alignment is widely used in phonetic research to align transcripts with acoustic signals. Yet there exists a lack of agreement on conventions for evaluating forced alignment, and our understanding of the reliability of forced aligners rests primarily on results from English. This study aims to fill these gaps by examining the concrete issue of forced aligning different Mandarin varieties. It evaluates machine-generated alignments from Montreal Forced Aligner (MFA); McAuliffe, Socolof, Mihuc, Wagner and Sonderegger Proc. Interspeech 2017, 498-502 (2017a) against two sets of independent human baselines using a Bayesian hierarchical multivariate regression model. Our findings suggest closer agreement between human aligners than between humans and MFA; large differences in alignment accuracy across different sequence types, with somewhat divergent patterns of errors across humans and MFA; some effects of speech rate and speaker-specific variation; and essentially no variation in robustness across varieties. These results serve (i) to reinforce previous results on the robustness of forced alignment across different varieties of the same language; and (ii) to provide a set of important methodological recommenations for evaluating forced alignment accuracy.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Liu et al. (2026) studied this question.

synapsesocial.com/papers/69db37f94fe01fead37c6072https://doi.org/10.1121/10.0043323
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Analysis of forced aligner performance on L2 English speech2024 · 4 citations
  2. 2The Mason-Alberta Phonetic Segmenter: a forced alignment system based on deep neural networks and interpolation2024 · 4 citations
  3. 3Fitting Linear Mixed-Effects Models Using lme42015 · 88,651 citations
  4. 4Polyglot and Speech Corpus Tools: A System for Representing, Integrating, and Querying Speech Corpora2017 · 17 citations
  5. 5Practical Bayesian model evaluation using leave-one-out cross-validation and WAIC2016 · 5,638 citations