Models based on Transformers have attained the highest performance benchmarks in neural machine translation, which benefits from the powerful learning ability of the model. However, its size prohibits its deployment on edge devices due to being excessively large, e.g., smartphones or embedded sensors. To address this issue, some model compression techniques are proposed, and knowledge distillation (KD) is a popular one. Previous research findings suggest that, when compared to the teacher model, the notable differences in the student model significantly influence the impact of KD. To alleviate the shortcoming, a multi-step knowledge distillation method is proposed to optimize the student model by filling the gap between the student and teacher models. Specifically, given a student (small) model and a teacher (large) model, a medium teacher-assistant model is introduced which is produced by the teacher model. Subsequently, the teacher-assistant model is employed to refine and Upgrade the student model’s functioning to achieve improved performance. Several experiments are performed across four translation tasks, and the findings from the experimental results underscore the effectiveness of the multi-step teacher-assistant KD method.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Improving Neural Machine Translation by Multi-step Teacher-Assistant Knowledge Distillation

  • Fuxue Li,
  • Haoming Ma,
  • Chuncheng Chi,
  • Hong Yan,
  • Beibei Liu

摘要

Models based on Transformers have attained the highest performance benchmarks in neural machine translation, which benefits from the powerful learning ability of the model. However, its size prohibits its deployment on edge devices due to being excessively large, e.g., smartphones or embedded sensors. To address this issue, some model compression techniques are proposed, and knowledge distillation (KD) is a popular one. Previous research findings suggest that, when compared to the teacher model, the notable differences in the student model significantly influence the impact of KD. To alleviate the shortcoming, a multi-step knowledge distillation method is proposed to optimize the student model by filling the gap between the student and teacher models. Specifically, given a student (small) model and a teacher (large) model, a medium teacher-assistant model is introduced which is produced by the teacher model. Subsequently, the teacher-assistant model is employed to refine and Upgrade the student model’s functioning to achieve improved performance. Several experiments are performed across four translation tasks, and the findings from the experimental results underscore the effectiveness of the multi-step teacher-assistant KD method.