PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 2, 2026Buana Information Technology and Computer Sciences (BIT and CS)0 citationsOpen Access

Implementation of the LSTM Model for Speech-to-Text Systems in the Recognition of the Walikan Language of Malang

View Full Paper
RRRaynanda RaynandaARAviv Yuniar RahmanIIIstiadi Istiadi

Key Points

  • The research aims to develop an effective Speech-to-Text system for the unique Malang Walikan language using the LSTM model.
  • Developed an STT system based on the LSTM model
  • Collected 1,000 sentences from social media and recordings
  • Processed data using Mel Frequency Cepstral Coefficients (MFCC)
  • Evaluated performance with Word Error Rate (WER), Character Error Rate (CER), and Average Test Loss metrics.
  • Achieved a WER value of 1.0 with a 699:300 data split
  • Found a CER of 0.78 with a 799:200 split
  • Recorded an Average Test Loss of 11.0147 with a 299:700 split
  • Identified challenges in model performance due to high Average Test Loss and potential overfitting.

Abstract

This study developed a Speech-to-Text (STT) system based on the Long Short-Term Memory (LSTM) model to recognize and convert speech in the Malang Walikan language into text. The Malang Walikan language has a unique linguistic structure in the form of word reversal, which poses a challenge in speech recognition. The data used consisted of 1,000 sentences collected from social media and direct recordings. The data was processed using Mel Frequency Cepstral Coefficients (MFCC) and then used to train the LSTM model.The system's performance was evaluated using the Word Error Rate (WER), Character Error Rate (CER), and Average Test Loss metrics. The best results obtained showed a WER value of 1.0 on a 699:300 data split, a CER of 0.78 on a 799:200 split, and an Average Test Loss of 11.0147 on a 299:700 split.The high Average Test Loss value indicates the model's difficulty in minimizing prediction errors, which may be caused by the model's mismatch with the data patterns or overfitting. To improve the model's performance, it is recommended to improve the quality of the training data, optimize the parameters, and apply regularization techniques.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Raynanda et al. (2026) studied this question.

synapsesocial.com/papers/6980ffc6c1c9540dea81291dhttps://doi.org/10.36805/m0pcpk09
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Development of a Deep Learning-Based Text-To-Speech System for the Malang Walikan Language Using the Pre-Trained SpeechT5 and Hifi-GAN Models2025
  2. 2Tongan Speech Recognition Based on Layer-Wise Fine-Tuning Transfer Learning and Lexicon Parameter Enhancement2025 · 1 citations
  3. 3Indonesian Language Sign Detection using Mediapipe with Long Short-Term Memory (LSTM) Algorithm2025 · 2 citations
  4. 4Javanese and Sundanese speech recognition using Whisper2025
  5. 5A Comparative Study of Khasi Speech Recognition Systems with Recurrent Neural Network-Based Language Model2024 · 1 citations