A novel unsupervised contrastive learning framework for ancient Yi script character dataset construction
摘要
The ancient Yi books have a long history and are one of the most important cultural heritages of humanity. Currently, Yi character image recognition datasets are all constructed by manual handwriting. Nevertheless, there are significant feature differences between handwritten character images and real ancient manuscript character images. This limits their effectiveness in practical ancient manuscript character recognition scenarios. To address these issues, we propose a novel dataset construction method to construct a real-world character dataset Yi_1047. Our method first uses an unsupervised contrastive learning model to perform accurate feature extraction. It applies this to a large number of unlabeled single-character images from real ancient manuscripts. Then, we use image retrieval and clustering methods to construct a real Yi script single-character dataset. It contains 1047 classes of ancient Yi script characters, totaling 265,094 character images. It is currently the only non-manually handwritten ancient Yi single-character dataset. Additionally, for highly similar characters, we specifically construct a comprehensive list of similar characters. Extensive experimental results demonstrate the effectiveness and superiority of our proposed method for constructing the Yi_1047 dataset. This innovative method not only achieves the digital preservation of ancient Yi script, but our experiments also show that it plays a role in constructing single-character recognition datasets for other ancient books.