PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 1, 2026ACM Transactions on Asian and Low-Resource Language Information Processing0 citations

Optimizing Low-Resource Machine Transliteration: A Case of Script Transition from Manipuri in Bengali Script to Meetei-Mayek

View Full Paper
GMGourashyam MoirangthemKNKishorjit Nongmeikapam

Key Points

  • The aim is to optimize machine transliteration for the Manipuri language in a low-resource context.
  • Compared three machine transliteration models: rule-based, statistical, and neural.
  • Designed a novel technique for building parallel datasets in low-resource settings.
  • Produced a gold-standard corpus of 35,000 transliterated Manipuri words.
  • Achieved a character error rate (CER) of 0.66.
  • Obtained a BLEU score of 98.7 and a METEOR score of 99.00.
  • The best encoder-decoder model set a new record for transliteration performance.

Abstract

At the expense of quantity and quality of training data, corpus-based models are becoming superior to rule-based models in solving complex Natural Language Processing (NLP) problems. In this research work, three categories of Machine Learning (ML) models for Machine Translatiteration (MTx) tasks are examined in a strictly low-resource scenario. This work studies and compares the Rule-Based Machine Transliteration (RBMTx) model, Statistical Machine Transliteration (SMTx) model and five generation-defining Neural Machine Transliteration (NMTx) models for the transliteration task in the low-resource Manipuri language. The work also discusses the contemporary script issues for the Manipuri language. The work explored and demonstrated how existing RBMTx models can facilitate corpus-based data-intensive machine learning models for low-resource languages using a novel technique for building parallel datasets. This study produced a gold-standard corpus of 35,000 Bengali script-Meetei Mayek parallel Manipuri words. With a Character Error Rate (CER) of only 0.66, a chrF score of 98.3, a BLEU score of 98.7 and a METEOR score of 99.00, the best performing Encoder-Decoder with Self Attention machine transliteration model sets a new performance record for the Bengali script to Meetei Mayek transliteration task. In addition to its immense potential for facilitating the ongoing script transition from Bengali script to Meetei Mayek, this research work will also help in addressing the low-resource bottleneck of Meetei Mayek for downstream Manipuri language NLP tasks.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Moirangthem et al. (2026) studied this question.

synapsesocial.com/papers/69cd7ae65652765b073a870fhttps://doi.org/10.1145/3806198
Ask AI
Helpful
Bookmark
Share
View Full Paper