Document Recognition and Analysis is a long-standing area of study that has been actively investigated for many decades. Handwritten text recognition (HTR) in Indian scripts is still in its early stages since the lack of publicly available datasets hampers the development of offline text recognition algorithms. Comparisons become tiresome due to the intricate structure of scripts, variances in depiction, and the presence of cursive writing. Hindi is the fourth most prevalent spoken language in the world, spoken by 615 million people, whereas Bengali is the sixth most popular, spoken by 265 million people [Source] . Both are read in a left-to-right direction. The accessible datasets of both languages are limited in size, have a limited number of writer’s samples, and use limited annotations, which poses challenges in developing resilient solutions utilising contemporary machine learning methods. Here, we have prepared Hindi and Bengali text scripts, each covering all the letters of Hindi and Bengali literature, named AIO-HB(All in one Hindi-Bengali Dataset). The dataset, which can be accessed at https://sites.google.com/view/aio-hb-dataset , has also been made public for the benefit of the researchers. These handwritten scripts have been written by 202 different writers multiple times. These scripts are considered from individuals of different professions, including diverse ages and genders. The AIO-HB dataset is benchmarked using conventional deep-learning models for Handwritten Text Recognition (HTR). It can be used for various document image analysis applications, such as recognising script sentences, segmenting text lines, segmenting words, detecting words, and identifying writers. This article explores the reasons for improvements in HTR performance across scripts and the utility of annotation for Indian HTRs.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

AIO-HB: A Handwritten Text Image Dataset of Hindi and Bengali Indian Scripts for Handwritten Text Recognition

  • Piyush Kanti Samanta,
  • Samit Biswas

摘要

Document Recognition and Analysis is a long-standing area of study that has been actively investigated for many decades. Handwritten text recognition (HTR) in Indian scripts is still in its early stages since the lack of publicly available datasets hampers the development of offline text recognition algorithms. Comparisons become tiresome due to the intricate structure of scripts, variances in depiction, and the presence of cursive writing. Hindi is the fourth most prevalent spoken language in the world, spoken by 615 million people, whereas Bengali is the sixth most popular, spoken by 265 million people [Source] . Both are read in a left-to-right direction. The accessible datasets of both languages are limited in size, have a limited number of writer’s samples, and use limited annotations, which poses challenges in developing resilient solutions utilising contemporary machine learning methods. Here, we have prepared Hindi and Bengali text scripts, each covering all the letters of Hindi and Bengali literature, named AIO-HB(All in one Hindi-Bengali Dataset). The dataset, which can be accessed at https://sites.google.com/view/aio-hb-dataset , has also been made public for the benefit of the researchers. These handwritten scripts have been written by 202 different writers multiple times. These scripts are considered from individuals of different professions, including diverse ages and genders. The AIO-HB dataset is benchmarked using conventional deep-learning models for Handwritten Text Recognition (HTR). It can be used for various document image analysis applications, such as recognising script sentences, segmenting text lines, segmenting words, detecting words, and identifying writers. This article explores the reasons for improvements in HTR performance across scripts and the utility of annotation for Indian HTRs.