Using MSA Data to Generate Training Data for Translating the Libyan Dialect into MSA Through Word Replacement and Fine-tuning Strategies.

Authors

  • Husien Alhammi University of Sfax

DOI:

https://doi.org/10.24996/ijs.2026.67.10.24

Keywords:

machine translation, MSA monolingual data, training data, word replacement, Libyan dialect, NMT model.

Abstract

The translation of dialectal Arabic into modern standard Arabic using neural machine translation approaches is a significant challenge due to the scarcity of parallel corpora. Low-resource languages such as the Libyan dialect often suffer from a lack of sufficient parallel training data. To address this issue, we introduce an innovative method for generating the Libyan dialect-modern standard Arabic training dataset from monolingual modern standard Arabic sentences using word replacement techniques. This research presents a method that combines word replacement with fine-tuning to build a neural machine translation model for translating the Libyan dialect into modern standard Arabic. The results show that using word replacement and fine-tuning techniques noticeably improves the translation accuracy compared to baselined systems.

Issue

Section

Computer Science

How to Cite

[1]
H. Alhammi, “Using MSA Data to Generate Training Data for Translating the Libyan Dialect into MSA Through Word Replacement and Fine-tuning Strategies”., Iraqi Journal of Science, vol. 67, no. 10, doi: 10.24996/ijs.2026.67.10.24.