This study focuses on the speech interaction experience of elderly users and proposes a speech recognition architecture, Con-DFSMN, which integrates a Convolution-augmented Transformer (Conformer) and Deep Feedforward Sequential Memory Network (DFSMN). The aim is to enhance recognition accuracy and system stability in variable and complex speech environments. The Additive Margin-Circle Loss (AM-Circle Loss) compound loss function is introduced, which effectively enhances the ability to distinguish class boundaries. This enables the model to exhibit excellent generalization performance when processing speech with high intra-class fluctuation and ambiguous boundaries. For the Tsinghua Chinese 30-hour Speech Corpus (THCHS-30) and DemantiaBank datasets, experimental results show that Con-DFSMN achieves accuracies of 95.3% and 91.8% in short speech recognition, respectively, representing a 13.1% improvement over traditional Transformer models. Con-DFSMN also maintains advantages in long speech recognition, with maximum accuracies of 90.8% and 87.3%. During training, Con-DFSMN demonstrates faster convergence speed and more stable loss decline, advancing 15 epochs earlier than Deepspeech2 and improving by approximately 25% compared to Transformer. The results verify the effectiveness and robustness of the model, demonstrating its adaptability to the speech features of the elderly, laying a foundation for the development of intelligent elderly speech interaction systems.
Zhang et al. (Sat,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: