The paper presents an analysis of Cyrillic Handwritten Text Recognition (HTR), high-lighting the unique challenges and technological advancements in this field across various countries. Despite the significance of HTR for digitizing historical documents and automating document processing, progress remains limited compared to Latin script, primarily due to the complexity of the Cyrillic character set and the variability in handwritten styles. The study underscores the achievements in Russian HTR, particularly through the adoption of advanced deep learning models, including CNNs, RNNs, and Transformer-based approaches. However, it identifies ongoing issues such as the scarcity of high-quality, labeled datasets and the need for more robust recognition systems. Comparative insights from other countries, including Kyrgyzstan, Uzbekistan, and additional Cyrillic-using nations such as Bulgaria, Serbia, North Macedonia, Belarus, Tajikistan, Ukraine, and Transnistria, reveal a common struggle with resource limitations and data availability. The paper advocates for increased investments and international collaboration to develop comprehensive datasets and enhance recognition technologies, ultimately aiming to improve the efficiency and accessibility of digital document processing for Cyrillic script users.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparative Analysis of Handwritten Text Recognition Across Cyrillic-Using Countries

  • Burul Shambetova,
  • Ruslan Isaev,
  • Nazira Abdillaeva,
  • Mekia Shigute Gaso

摘要

The paper presents an analysis of Cyrillic Handwritten Text Recognition (HTR), high-lighting the unique challenges and technological advancements in this field across various countries. Despite the significance of HTR for digitizing historical documents and automating document processing, progress remains limited compared to Latin script, primarily due to the complexity of the Cyrillic character set and the variability in handwritten styles. The study underscores the achievements in Russian HTR, particularly through the adoption of advanced deep learning models, including CNNs, RNNs, and Transformer-based approaches. However, it identifies ongoing issues such as the scarcity of high-quality, labeled datasets and the need for more robust recognition systems. Comparative insights from other countries, including Kyrgyzstan, Uzbekistan, and additional Cyrillic-using nations such as Bulgaria, Serbia, North Macedonia, Belarus, Tajikistan, Ukraine, and Transnistria, reveal a common struggle with resource limitations and data availability. The paper advocates for increased investments and international collaboration to develop comprehensive datasets and enhance recognition technologies, ultimately aiming to improve the efficiency and accessibility of digital document processing for Cyrillic script users.