Due to the severe scarcity of training data and the challenges in compiling dictionaries for Todo Mongolian, there have been no text recognition tools specifically designed for Todo Mongolian to date. This study aims to fill the research gap in Optical Character Recognition (OCR) for Todo Mongolian and introduces the first publicly available Todo Mongolian OCR dataset. We have developed a novel segmentation-free recognition method for entire lines of Todo Mongolian text, which draws on traditional Mongolian text recognition techniques. Data synthesis and enhancement techniques were used to expand the dataset and alleviate the issues of data scarcity. As part of the research, we have compiled and released a database containing 150,000 lines of generated Todo Mongolian text images, each meticulously annotated. The database, along with the scripts used for generating synthetic images and the data generation code, will be made freely available to the academic community to support further research. This method has been experimentally validated and achieved a word-level error rate of 15.27%. This work not only provides an initial solution to the OCR challenges of Todo Mongolian but also offers a valuable data resource for researchers in the field.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Segmentation-Free Todo Mongolian OCR and its Public Dataset

  • Weiqi Wang,
  • Feilong Bao,
  • Hui Zhang

摘要

Due to the severe scarcity of training data and the challenges in compiling dictionaries for Todo Mongolian, there have been no text recognition tools specifically designed for Todo Mongolian to date. This study aims to fill the research gap in Optical Character Recognition (OCR) for Todo Mongolian and introduces the first publicly available Todo Mongolian OCR dataset. We have developed a novel segmentation-free recognition method for entire lines of Todo Mongolian text, which draws on traditional Mongolian text recognition techniques. Data synthesis and enhancement techniques were used to expand the dataset and alleviate the issues of data scarcity. As part of the research, we have compiled and released a database containing 150,000 lines of generated Todo Mongolian text images, each meticulously annotated. The database, along with the scripts used for generating synthetic images and the data generation code, will be made freely available to the academic community to support further research. This method has been experimentally validated and achieved a word-level error rate of 15.27%. This work not only provides an initial solution to the OCR challenges of Todo Mongolian but also offers a valuable data resource for researchers in the field.