Classification of handwritten characters is a technique of optical character recognition, that helps us to recognize and process non-digital text found in our everyday surroundings especially in academic institutions. This paper; with a view to solving handwritten character recognition; presents a method to classify ‘Bengali’ characters. Bengali has been chosen for two main reasons, firstly, population-wise it is 6th most used languages in the world [1], secondly, there’s scarcity of reliable services where Bengali can be used effortlessly [2]. Successful recognition of these characters can lead us to achieve fast digitalization of text data and can widen the usage of optical character recognition-based services among a huge number of people. It can also later assist in language interchanging scenarios, like automated translation. The core of the research is to develop convolutional neural network-based classification model and to compare with the current available methods of Bengali character recognition. The experiments and evaluations are done on ‘Banglalekha Isolated dataset’. This dataset contains preprocessed binary images of the characters. We tried to find an additional preprocessing pipeline; comprised of histogram of oriented gradients, morphological transform, ROI (region of interest) extraction. It has yielded satisfactory accuracy in classifying 84 classes of characters on ‘Banglalekha Isolated dataset’.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Study on the Comparison of Various Preprocessing Methods for Bengali Character Recognition Using Convolutional Neural Network

  • Md Shafayet Jamil,
  • Hiroki Tamura

摘要

Classification of handwritten characters is a technique of optical character recognition, that helps us to recognize and process non-digital text found in our everyday surroundings especially in academic institutions. This paper; with a view to solving handwritten character recognition; presents a method to classify ‘Bengali’ characters. Bengali has been chosen for two main reasons, firstly, population-wise it is 6th most used languages in the world [1], secondly, there’s scarcity of reliable services where Bengali can be used effortlessly [2]. Successful recognition of these characters can lead us to achieve fast digitalization of text data and can widen the usage of optical character recognition-based services among a huge number of people. It can also later assist in language interchanging scenarios, like automated translation. The core of the research is to develop convolutional neural network-based classification model and to compare with the current available methods of Bengali character recognition. The experiments and evaluations are done on ‘Banglalekha Isolated dataset’. This dataset contains preprocessed binary images of the characters. We tried to find an additional preprocessing pipeline; comprised of histogram of oriented gradients, morphological transform, ROI (region of interest) extraction. It has yielded satisfactory accuracy in classifying 84 classes of characters on ‘Banglalekha Isolated dataset’.