Despite the prevalence of subword tokenization, its deterministic nature—splitting words into unique output tokens, may limit models from fully exploiting the intricate semantic compositions within words. Subword regularization methods address this limitation by using multiple subword sequences generated by tokenization. However, existing methods ineffectively utilize multi-granularity semantic compositions inherent in words, which is crucial for language understanding. In this paper, we propose Progressive and Consistent Subword Regularization (PCSR), a novel and simple subword regularization method that progressively changes the granularity of tokenization from fine to coarse dynamically during training and enforces the consistency between these multi-granularity subword segmentations. Moreover, we verify empirically that applying consistency constraints to existing subword regularization methods significantly improves their effectiveness for neural machine translation (NMT). Experiments on IWSLT and WMT translation tasks show that PCSR outperforms various subword regularization methods and their combinations, with BLEU score improvements up to 2.4 over the standard BPE baseline.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Progressive and Consistent Subword Regularization for Neural Machine Translation

  • Yongqi Gao,
  • Yingfeng Luo,
  • Qinghong Zhang,
  • Huibo Shao,
  • Tong Xiao,
  • Jingbo Zhu

摘要

Despite the prevalence of subword tokenization, its deterministic nature—splitting words into unique output tokens, may limit models from fully exploiting the intricate semantic compositions within words. Subword regularization methods address this limitation by using multiple subword sequences generated by tokenization. However, existing methods ineffectively utilize multi-granularity semantic compositions inherent in words, which is crucial for language understanding. In this paper, we propose Progressive and Consistent Subword Regularization (PCSR), a novel and simple subword regularization method that progressively changes the granularity of tokenization from fine to coarse dynamically during training and enforces the consistency between these multi-granularity subword segmentations. Moreover, we verify empirically that applying consistency constraints to existing subword regularization methods significantly improves their effectiveness for neural machine translation (NMT). Experiments on IWSLT and WMT translation tasks show that PCSR outperforms various subword regularization methods and their combinations, with BLEU score improvements up to 2.4 over the standard BPE baseline.