Spelling error correction (SEC) is an important but challenging task, serving as a foundation for various downstream tasks. Vietnamese spelling error correction confronts two main challenges: 1) annotated data is insufficient, and 2) state-of-the-art method struggles to correct out-of-vocabulary (OOV) words and spacing errors. To address these challenges, we propose an effective method for constructing high-quality training data that includes three data construction strategies. Additionally, we introduce a two-stage character-level generative spelling error correction framework (TS-CGSEC). This framework enhances the correction of OOV words and spacing errors by precisely modifying detected erroneous syllables. Experiments on both the VSEC and Viwiki-Spelling datasets demonstrate the effectiveness of the proposed training data construction method and the TS-CGSEC framework.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Character-Level Generative Network for Vietnamese Spelling Error Correction

  • Yanmei Ou,
  • Zhuofan You,
  • Siyi Lian,
  • Cuiyi Zhi,
  • Shengyi Jiang

摘要

Spelling error correction (SEC) is an important but challenging task, serving as a foundation for various downstream tasks. Vietnamese spelling error correction confronts two main challenges: 1) annotated data is insufficient, and 2) state-of-the-art method struggles to correct out-of-vocabulary (OOV) words and spacing errors. To address these challenges, we propose an effective method for constructing high-quality training data that includes three data construction strategies. Additionally, we introduce a two-stage character-level generative spelling error correction framework (TS-CGSEC). This framework enhances the correction of OOV words and spacing errors by precisely modifying detected erroneous syllables. Experiments on both the VSEC and Viwiki-Spelling datasets demonstrate the effectiveness of the proposed training data construction method and the TS-CGSEC framework.