A new evaluation method: evaluation data and metrics for Chinese grammatical error correction
摘要
As a fundamental task in natural language processing (NLP), Chinese Grammatical Error Correction (CGEC; Zhao et al., 2019; Tang et al., 2021; Zhao and Wang, 2020) has gradually received widespread attention and become a research hotspot. However, one obvious deficiency of the existing CGEC evaluation systems is that the evaluation values of the same error correction models are significantly influenced by the Chinese word segmentation (CWS) results or different language models. However, it is expected that these metrics should be independent of the CWS results and language models for a fair evaluation. To this end, we propose three novel evaluation metrics for CGEC in two dimensions: reference-based and reference-less. What’s more, according to these three evaluation metrics, we build a new evaluation metric that can comprehensively evaluate the CGEC model from multiple dimensions. We deeply evaluate and analyze the reasonableness and validity of the proposed metrics, and we expect them to become a new standard for CGEC.