Traditional ASR systems often struggle with low-resource languages, data dependence, and robustness in noisy or domain-specific environments. Telemedicine, clinical recordkeeping, and multilingual communication depend on accurate speech-to-text systems. This work introduces Wav2Vec-RCNet, a novel ASR architecture that integrates self-supervised Wav2Vec representations with recurrent neural networks (RNNs) and combines Cuckoo Search-tuned RNNs. Contrastive predictive coding lets Wav2Vec generate high-dimensional latent embeddings from raw audio. This is possible because effective feature representations do not require large annotated corpora. LSTM and GRU embeddings capture the temporal and sequential connections needed for accurate voice transcription. The Lévy flight-based search dynamics of Cuckoo Search optimise learning rates, hidden unit dimensions, dropout probabilities, and attention weight distributions for optimum performance. Wav2Vec-RCNet outperforms clinical benchmark datasets, multi-accent datasets, and conventional supervised and hybrid ASR systems in terms of word and character error rates. Wav2Vec-RCNet constructs a scalable, resource-efficient, and domain-adaptive framework for real-world speech-to-text applications, leveraging self-supervised feature learning, recurrent sequence modeling, and metaheuristic-driven optimization.
AlTaei et al. (Thu,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: