Character-Level Generative Network for Vietnamese Spelling Error Correction
摘要
Spelling error correction (SEC) is an important but challenging task, serving as a foundation for various downstream tasks. Vietnamese spelling error correction confronts two main challenges: 1) annotated data is insufficient, and 2) state-of-the-art method struggles to correct out-of-vocabulary (OOV) words and spacing errors. To address these challenges, we propose an effective method for constructing high-quality training data that includes three data construction strategies. Additionally, we introduce a two-stage character-level generative spelling error correction framework (TS-CGSEC). This framework enhances the correction of OOV words and spacing errors by precisely modifying detected erroneous syllables. Experiments on both the VSEC and Viwiki-Spelling datasets demonstrate the effectiveness of the proposed training data construction method and the TS-CGSEC framework.