PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 23, 2025Computers3 citationsOpen Access

Benchmarking the Responsiveness of Open-Source Text-to-Speech Systems

View Full Paper
HDHa Pham Thien DinhRPRutherford Agbeshi PatamiaMLMing Liu

Key Points

  • Responsiveness is crucial for real-time applications, unlike typical quality metrics for speech systems.
  • The benchmark reveals variability in latency, with some TTS models offering near-instant output while others lag.
  • Our open-source framework incorporates intelligibility and latency metrics for accurate, reproducible comparisons.
  • This groundwork supports further comprehensive assessments of TTS systems, combining responsiveness with quality.

Abstract

Responsiveness—the speed at which a text-to-speech (TTS) system produces audible output—is critical for real-time voice assistants yet has received far less attention than perceptual quality metrics. Existing evaluations often touch on latency but do not establish reproducible, open-source standards that capture responsiveness as a first-class dimension. This work introduces a baseline benchmark designed to fill that gap. Our framework unifies latency distribution, tail latency, and intelligibility within a transparent and dataset-diverse pipeline, enabling a fair and replicable comparison across 13 widely used open-source TTS models. By grounding evaluation in structured input sets ranging from single words to sentence-length utterances and adopting a methodology inspired by standardized inference benchmarks, we capture both typical and worst-case user experiences. Unlike prior studies that emphasize closed or proprietary systems, our focus is on establishing open, reproducible baselines rather than ranking against commercial references. The results reveal substantial variability across architectures, with some models delivering near-instant responses while others fail to meet interactive thresholds. By centering evaluation on responsiveness and reproducibility, this study provides an infrastructural foundation for benchmarking TTS systems and lays the groundwork for more comprehensive assessments that integrate both fidelity and speed.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Dinh et al. (2025) studied this question.

synapsesocial.com/papers/68d4759931b076d99fa6d85dhttps://doi.org/10.3390/computers14100406
Ask AI
Helpful
Bookmark
Share
View Full Paper