Videoconference systems have become indispensable across generations and professional domains. In business settings, the need for immediate archiving and summarization of online meetings has driven the widespread adoption of AI-based transcription and summarization services. However, speech degradation caused by unstable network conditions or system limitations significantly impairs transcription accuracy. While listeners can often perceive when speech quality has degraded, identifying and addressing the root cause is challenging for the speaker experiencing the issue, particularly when the degradation stems from poor network performance. This study proposes the development of a speech quality evaluation model designed to assess transcription accuracy under conditions of packet loss. Network-induced speech degradation is simulated using the SILK audio codec, which is widely employed in modern videoconference systems. By analyzing the relationship between degraded speech and transcription performance, the model aims to objectively evaluate whether the audio quality of a videoconference is sufficient for accurate transcription. The proposed model is expected to serve as a practical tool for improving audio quality preparation, thereby enabling more accurate transcriptions and efficient post-meeting documentation.
Takeuchi et al. (Wed,) studied this question.