Machine translation from Arabic dialects to modern standard Arabic has recently gained attention. The quality of machine translation remains a challenge, and additional work is needed for improvement. This research presents an automatic machine translation system for translating the Libyan dialect short dual-form sentences to Modern Standard Arabic, which is the first machine translation effort to translate the Libyan dialect to modern standard Arabic. However, our work involves leveraging a pre-existing Libyan dialect resource, a bilingual dictionary, to build a machine translation system. The system uses a rule-based approach that relies on explicitly defined linguistic rules for the Libyan dialect and modern standard Arabic. In this paper, the translation process is divided into four phases: analysis, lexical transfer, lexical selection, and generation. In addition, to produce the fluent modern standard Arabic sentences, a 2-gram parallel corpus was used. The evaluation process is done by comparing the system results with the original human translation. The system was assessed using a test dataset of 100 short dual-form sentences. The BLEU, TER and ChrF scores of the system were 41.37%, 40.3% and 66.9%, respectively.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Rule-Based System for Translating Libyan Dialect Dual Forms to Modern Standard Arabic

  • Husien Alhammi,
  • Kais Haddar

摘要

Machine translation from Arabic dialects to modern standard Arabic has recently gained attention. The quality of machine translation remains a challenge, and additional work is needed for improvement. This research presents an automatic machine translation system for translating the Libyan dialect short dual-form sentences to Modern Standard Arabic, which is the first machine translation effort to translate the Libyan dialect to modern standard Arabic. However, our work involves leveraging a pre-existing Libyan dialect resource, a bilingual dictionary, to build a machine translation system. The system uses a rule-based approach that relies on explicitly defined linguistic rules for the Libyan dialect and modern standard Arabic. In this paper, the translation process is divided into four phases: analysis, lexical transfer, lexical selection, and generation. In addition, to produce the fluent modern standard Arabic sentences, a 2-gram parallel corpus was used. The evaluation process is done by comparing the system results with the original human translation. The system was assessed using a test dataset of 100 short dual-form sentences. The BLEU, TER and ChrF scores of the system were 41.37%, 40.3% and 66.9%, respectively.