In recent years, while significant progress has been made in speech recognition technology for high-resource languages (such as English and Mandarin), research on low-resource languages with complex phonologies, like Tibetan, has progressed relatively slow. As a low-resource and phonologically complex language, Amdo Tibetan faces dual challenges in speech recognition: data scarcity and insufficient quality and diversity of available datasets. The lack of publicly accessible datasets has imposed numerous constraints on related research. To address these challenges, this paper introduces and presents an open-source speech recognition dataset for the Amdo Tibetan dialect. The speech samples were initially collected in Xiahe County, Gansu Province, China, comprising 31 hours of recordings from 66 native speakers along with corresponding transcriptions. Subsequent manual quality control and standardization were applied to ensure the authenticity of the dialect as well as the consistency and quality of the data. All resources in this dataset have been made publicly available and have already been utilized in multiple research papers and studies on Tibetan speech recognition, receiving widespread acclaim from experts in the field—further validating the dataset’s quality. This dataset serves as an important supplement to high-quality speech data for Amdo Tibetan, and provides unique support for cross-lingual transfer learning and few-shot speech technology research due to its complex phonological characteristics.
XIE et al. (Sun,) studied this question.