Mixed Multimodal Contrastive Learning for Enhancing Detection and Correction of Faked and Misspelled Chinese Characters
摘要
Visual Chinese Character Checking (C3) aims to detect and correct errors in handwritten Chinese text images, including faked characters and misspelled characters. This task is beneficial for subsequent tasks by improving the efficiency of identifying errors in handwritten text. Recent methods are mainly based on Optical Character Recognition (OCR) and Pre-trained Language Models (PLMs). Visual Chinese Character Checking is an emerging task, and relevant research has made progress. However, we believe that existing work has not fully leveraged the inherent knowledge of pre-trained models and has not addressed the semantic bias issue between pre-trained models and the character checking task. These challenges result in deficiencies in recognizing misspelled Chinese characters and correcting misused characters. Therefore, we propose various multimodal contrastive learning methods based on image-to-image and image-to-text comparisons. These methods are used throughout the processes of character recognition, error detection, and correction. By aligning the semantic feature representations among different models, our approach makes these models more suitable for the Visual Chinese Character Checking task, thereby enhancing their capabilities.