Using MSA Data to Generate Training Data for Translating the Libyan Dialect into MSA Through Word Replacement and Fine-tuning Strategies.
DOI:
https://doi.org/10.24996/ijs.2026.67.10.24Keywords:
machine translation, MSA monolingual data, training data, word replacement, Libyan dialect, NMT model.Abstract
The translation of dialectal Arabic into modern standard Arabic using neural machine translation approaches is a significant challenge due to the scarcity of parallel corpora. Low-resource languages such as the Libyan dialect often suffer from a lack of sufficient parallel training data. To address this issue, we introduce an innovative method for generating the Libyan dialect-modern standard Arabic training dataset from monolingual modern standard Arabic sentences using word replacement techniques. This research presents a method that combines word replacement with fine-tuning to build a neural machine translation model for translating the Libyan dialect into modern standard Arabic. The results show that using word replacement and fine-tuning techniques noticeably improves the translation accuracy compared to baselined systems.




