PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 31, 2026Scientific Reports0 citationsOpen Access

Reinforcement learning from human feedback improves automatic speech recognition models for ancient Thai language

JPJettasic PopunWLWilaiporn LeeKSKanabadee Srisomboon

Key Points

  • The study aims to improve automatic speech recognition for ancient Thai Traditional Medicine recitations using reinforcement learning from human feedback.
  • Developed an ASR framework enhanced with RLHF to transcribe TTM recitations.
  • Created a domain-specific corpus expanded from 2,000 to 14,000 utterances using data augmentation techniques like pitch shifting and time stretching.
  • Evaluated Whisper-small and Wav2Vec2 architectures before and after optimization using a reward model.
  • Whisper-small achieved a 42.2% reduction in Word Error Rate (WER).
  • Demonstrated superior tonal preservation over supervised baselines.

Abstract

Thailand’s ancient medicinal manuscripts preserve centuries of therapeutic wisdom but are threatened by physical deterioration. Conventional digitization via Optical Character Recognition (OCR) is largely ineffective due to faded ink and archaic orthography on palm leaves. Consequently, oral recitation by expert practitioners remains the most reliable method for preserving this content. This study proposes an Automatic Speech Recognition (ASR) framework enhanced with Reinforcement Learning from Human Feedback (RLHF) to transcribe ancient Thai Traditional Medicine (TTM) recitations. We developed a domain-specific corpus and applied rigorous data augmentation techniques (pitch shifting and time stretching) to expand the dataset from 2,000 to 14,000 utterances, recorded from ten certified practitioners. Two state-of-the-art architectures, Whisper-small and Wav2Vec2, were evaluated before and after optimization using a reward model trained on expert linguistic judgments of phonetic, tonal, and semantic fidelity. Experimental results demonstrate that RLHF, combined with the augmented dataset, substantially improves transcription quality. Whisper-small achieved a 42.2% reduction in Word Error Rate (WER) and demonstrated superior tonal preservation compared to supervised baselines. These findings highlight the effectiveness of human-aligned ASR for low-resource tonal languages and support the scalable digitization of endangered medical heritage.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Popun et al. (2026) studied this question.

synapsesocial.com/papers/6a1bd12d5783ba022b6fcc77https://doi.org/10.1038/s41598-026-54756-x
Ask AI
Helpful
Bookmark
Share
View Full Paper