Proposed method of acquiring train data for early-modern Japanese printed character recognizers
摘要
The National Diet Library maintains a digital archive of approximately 400,000 historical texts from the Meiji to early Showa periods (1868–1935), representing a significant corpus of Japanese cultural heritage. While these materials are digitally preserved as high-resolution images, the lack of machine-readable text severely constrains their accessibility and scholarly utility. This limitation particularly impacts early-modern Japanese printed texts, where conventional optical character recognition systems achieve only 60–75% accuracy due to the unique characteristics of historical typography. Current approaches to digitizing these materials face significant challenges: manual transcription is prohibitively expensive for large-scale application, while automated methods struggle with the irregular features of early-modern printed characters, including non-standardized character forms, ink bleeding, and letterpress artifacts. Furthermore, the scarcity of available character samples, particularly for less frequent characters, hinders the development of robust recognition systems. This study presents a novel methodology for enhancing early-modern Japanese character recognition through a two-fold approach: (1) utilizing CycleGAN-generated synthetic characters to augment limited historical samples and (2) developing an optimal mixing strategy between original, generated, and modern characters. Our approach achieves recognition rates of up to 97.81%, significantly outperforming conventional methods while addressing the critical challenge of data scarcity in historical character recognition. Experimental validation demonstrated that our proposed method achieved substantial improvements in recognition accuracy: incorporating generated characters increased recognition rates from 93.43 to 97.81% (p < 0.001) when sufficient original characters were available, and maintained robust performance (94.90%) even with limited historical samples. The model achieved AUC scores consistently above 0.999 across all experiments, demonstrating high reliability in character discrimination regardless of the data composition.