Utilizing a small amount of domain knowledge to achieve high-quality domain-specific translation is a challenging task. Nowadays, Large language model(LLM) is capable of generating more fluent and human-preferred translations through personalized instructions. However, in the field of domain-specific machine translation, the performance of LLM is inferior to traditional methods due to the lack of domain training data and the absence of domain transfer ability. To address the issue, we decided to incorporate terminology knowledge, which is crucial for accurately capturing the precise semantics of domain-specific texts. We design two types of terminology alignment instructions to enhance the model’s cross-linguistic terminology alignment capability, explicitly integrating terminology knowledge into the model training process. According to the experiment, the model fine-tuned with MT+G-Align significantly outperformed the baseline through terminology translation accuracy and translation quality, demonstrating the effectiveness of the terminology alignment instructions. On the WMT 2023 Terminology Translation task, experimental results show that our approach achieves the best results in all three directions, including German-to-English, Chinese-to-English, and English-to-Chinese.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Incorporating Terminology Knowledge into Large Language Model for Domain-Specific Machine Translation

  • Xuan Zhao,
  • Chong Feng,
  • Shuanghong Huang,
  • Jiangyu Wang,
  • Haojie Xu

摘要

Utilizing a small amount of domain knowledge to achieve high-quality domain-specific translation is a challenging task. Nowadays, Large language model(LLM) is capable of generating more fluent and human-preferred translations through personalized instructions. However, in the field of domain-specific machine translation, the performance of LLM is inferior to traditional methods due to the lack of domain training data and the absence of domain transfer ability. To address the issue, we decided to incorporate terminology knowledge, which is crucial for accurately capturing the precise semantics of domain-specific texts. We design two types of terminology alignment instructions to enhance the model’s cross-linguistic terminology alignment capability, explicitly integrating terminology knowledge into the model training process. According to the experiment, the model fine-tuned with MT+G-Align significantly outperformed the baseline through terminology translation accuracy and translation quality, demonstrating the effectiveness of the terminology alignment instructions. On the WMT 2023 Terminology Translation task, experimental results show that our approach achieves the best results in all three directions, including German-to-English, Chinese-to-English, and English-to-Chinese.