This study aimed to evaluate the accuracy of the multilingual automatic speech recognition (ASR) model Whisper large-v3 (OpenAI, San Francisco, CA, USA) in recognizing isolated Korean monosyllables (consonant-vowel and consonant-vowel-consonant structures) produced by two older adults with sensorineural hearing loss, using expert audiologists as the perceptual reference standard. Two older Korean adults (87-year-old male, 78-year-old female) each produced 130 Korean monosyllables. Four audiologists independently transcribed the tokens under blinded conditions. The same audio files were processed using Whisper large-v3 (OpenAI) in an offline environment with Korean language settings. Phoneme-level accuracy, inter-rater agreement, and error patterns were compared across raters, speakers, and phoneme categories. Audiologist transcription showed moderate agreement (mean accuracy: 59.7% for S-001 and 61.7% for S-002; perfect agreement: 55.4% and 62.0%, respectively), reflecting variability in elderly speech perception. In contrast, Whisper demonstrated substantially lower accuracy (13.1% and 13.8%), with high error rates. Errors were primarily observed in specific phoneme categories, including tense consonants, fricatives, liquids, and mid-front vowels, whereas nasals and high vowels were relatively preserved. Whisper did not achieve higher transcription accuracy than audiologists for any monosyllable tested. Whisper large-v3 (OpenAI) shows substantially lower performance than human audiologists in phoneme-level transcription of Korean monosyllables produced by older adults with sensorineural hearing loss. These findings suggest that general-purpose ASR systems have limited clinical applicability in this population and should currently be considered as assistive tools rather than replacements for expert evaluation. Further development, including elderly- and Korean-specific model adaptation, is required.
Han et al. (Thu,) studied this question.