Plm-pos sponsored transformer-based ancient Chinese–English neural machine translation
摘要
Ancient Chinese (AC) texts are invaluable resources for studying ancient Chinese history and humanities. Automatic translation from AC to English enables international scholars to gain exposure to and understand the history of ancient China more effectively. Adding well-pre-trained contextualized representations from pre-trained language models (PLMs) and linguistic information such as Part-Of-Speech (POS) tags have been shown to help enhance Transformer-based neural machine translation (NMT) tasks. However, existing Transformer-based NMT methods either focus exclusively on PLM representations to improve translation quality or depend primarily on linguistic features. These approaches lead to an inability to fully utilize the potential of both PLM representations and language information. This paper proposes a new Pre-trained Language Model and Part-of-Speech (PLM-POS) sponsored Transformer-based ancient Chinese-English NMT model to enhance the quality of NMT with both PLM representations and POS dense vectors. The proposed model can decompose into two steps: First, Cloning is performed on the top m layers of the PLM, fine-tuning the cloned layers through supervised learning to obtain POS dense vectors while keeping the original PLM unchanged. Second, integrating the PLM representations and POS vectors into the Transformer’s Encoder and Decoder via attention mechanisms. The model can easily incorporate additional cloned layers to include more linguistic/non-linguistic information. Extensive experiments on AC-English, NOS AC-MC, Guwen-UNILM, and WMT18 datasets demonstrate the framework’s effectiveness, achieving improvements of 0.04