In recent decades, machine translation (MT) has undergone significant progress. The process of translating between natural languages is approached as a machine learning challenge in statistical machine translation (SMT). This method relies on training models on parallel corpora, such as the Dogri and English languages used in our experiment. In the expansive realm of machine translation, considerable efforts have been dedicated to exploring and advancing language pairs across the linguistic spectrum. Numerous languages have undergone extensive research and development, contributing to the evolution of machine translation technologies. However, amidst this wealth of exploration, the Dogri language has remained relatively unexplored in the context of machine translation. This study proposes a statistical machine translation system for the historically significant language of Dogri, which is listed in the Indian Constitution and recently designated as the official language of erstwhile state of Jammu and Kashmir. Dogri is a low-resource language as besides various fundamental tools for machine translation line Dogri-English dictionary and Dogri-English corpora do not exist. Author has manually generated a corpus of 10,305 sentences for the Dogri-English language pair for the study. In addition, a manually generated Dogri-English parallel corpus for developing SMT for Dogri containing 10,305 sentences employed for training, validating, and testing the SMT system. Statistical machine translation system for the Dogri-English language pair has been attempted for the first time, and the outcomes are optimistic for continued system improvement. Dogri-English parallel corpus when divided in the ratio 90:05:05, 80:10:10, and 70:15:15 for train set, validate set, and test set resulted in a BLEU score of 24.39, 22.44, and 21.69, respectively.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Advancing Dogri-English Translation Through Statistical Machine Translation Technique

  • Vijay Singh Sen,
  • Shubhnandan S. Jamwal

摘要

In recent decades, machine translation (MT) has undergone significant progress. The process of translating between natural languages is approached as a machine learning challenge in statistical machine translation (SMT). This method relies on training models on parallel corpora, such as the Dogri and English languages used in our experiment. In the expansive realm of machine translation, considerable efforts have been dedicated to exploring and advancing language pairs across the linguistic spectrum. Numerous languages have undergone extensive research and development, contributing to the evolution of machine translation technologies. However, amidst this wealth of exploration, the Dogri language has remained relatively unexplored in the context of machine translation. This study proposes a statistical machine translation system for the historically significant language of Dogri, which is listed in the Indian Constitution and recently designated as the official language of erstwhile state of Jammu and Kashmir. Dogri is a low-resource language as besides various fundamental tools for machine translation line Dogri-English dictionary and Dogri-English corpora do not exist. Author has manually generated a corpus of 10,305 sentences for the Dogri-English language pair for the study. In addition, a manually generated Dogri-English parallel corpus for developing SMT for Dogri containing 10,305 sentences employed for training, validating, and testing the SMT system. Statistical machine translation system for the Dogri-English language pair has been attempted for the first time, and the outcomes are optimistic for continued system improvement. Dogri-English parallel corpus when divided in the ratio 90:05:05, 80:10:10, and 70:15:15 for train set, validate set, and test set resulted in a BLEU score of 24.39, 22.44, and 21.69, respectively.