<p>In the context of neural machine translation (NMT) for spoken languages, the availability of bilingual data is often limited, leading to performance bottlenecks. Traditional approaches, such as morphological segmentation and byte-pair encoding, have shown promise but can struggle with informal texts and lack linguistic depth. To address these challenges, we propose a novel morphological segmentation method for spoken language NMT that leverages both monolingual and bilingual features. Our approach incorporates several monolingual features, including morphological segmentation, byte-pair encoding (BPE), and language model features, as well as bilingual features such as bilingual word alignment. We then use a log-linear model to combine these features and predict the final morphological segmentation results. Our experiments demonstrate that the morphemes generated by our proposed model significantly improve spoken language NMT performance. Specifically, we observe improvements of up to 1.9 BLEU points across several language pairs, including Uyghur-Chinese, Turkish-English, and Kazakh-English. These results suggest that our method is effective in enhancing NMT for spoken languages, particularly in low-resource scenarios.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Bilingual and monolingual features enhanced morphological segmentation for spoken language neural machine translation

  • Chenggang Mi,
  • Shaoliang Xie,
  • Qi Liu

摘要

In the context of neural machine translation (NMT) for spoken languages, the availability of bilingual data is often limited, leading to performance bottlenecks. Traditional approaches, such as morphological segmentation and byte-pair encoding, have shown promise but can struggle with informal texts and lack linguistic depth. To address these challenges, we propose a novel morphological segmentation method for spoken language NMT that leverages both monolingual and bilingual features. Our approach incorporates several monolingual features, including morphological segmentation, byte-pair encoding (BPE), and language model features, as well as bilingual features such as bilingual word alignment. We then use a log-linear model to combine these features and predict the final morphological segmentation results. Our experiments demonstrate that the morphemes generated by our proposed model significantly improve spoken language NMT performance. Specifically, we observe improvements of up to 1.9 BLEU points across several language pairs, including Uyghur-Chinese, Turkish-English, and Kazakh-English. These results suggest that our method is effective in enhancing NMT for spoken languages, particularly in low-resource scenarios.