<p>Recent, efforts to digitize books have aimed to enhance accessibility and convenience. However, early-modern Japanese books published between the Meiji and early Showa period (1868–1935) differ significantly from modern publications, posing challenges for character recognition systems. Although progress has been made in developing recognition systems tailored to these texts, recognition errors remain common. This paper proposes a method for detecting and correcting misrecognized characters in early-modern Japanese books using Bidirectional Encoder Representations from Transformers (BERT). The approach consists of three stages: erroneous sentence detection, recognition error identification, and error correction. Experimental results demonstrate that the proposed method achieves <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(94.7\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>94.7</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation> accuracy in detecting erroneous sentences and <InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(94.3\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>94.3</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation> accuracy in identifying recognition errors. Moreover, the correct character appeared among the top five correction candidates in <InlineEquation ID="IEq3"> <EquationSource Format="TEX">\(91.0\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>91.0</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation> of the identified errors.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Correction of textual errors in early-modern Japanese books

  • Hiyori Kanetaka,
  • Miho Chiyonobu,
  • Yuki Takemoto,
  • Yu Ishikawa,
  • Masami Takata

摘要

Recent, efforts to digitize books have aimed to enhance accessibility and convenience. However, early-modern Japanese books published between the Meiji and early Showa period (1868–1935) differ significantly from modern publications, posing challenges for character recognition systems. Although progress has been made in developing recognition systems tailored to these texts, recognition errors remain common. This paper proposes a method for detecting and correcting misrecognized characters in early-modern Japanese books using Bidirectional Encoder Representations from Transformers (BERT). The approach consists of three stages: erroneous sentence detection, recognition error identification, and error correction. Experimental results demonstrate that the proposed method achieves \(94.7\%\) 94.7 % accuracy in detecting erroneous sentences and \(94.3\%\) 94.3 % accuracy in identifying recognition errors. Moreover, the correct character appeared among the top five correction candidates in \(91.0\%\) 91.0 % of the identified errors.