The past few years have been marked by considerable growth of interest in handwritten recognition systems due to their usefulness in tasks such as automatic form processing, document scanning, and human–computer interface. The current paper offers an improved technique for text and handwriting digit identification through the incorporation of CNNs, OCR, and language translation technology. The aim of this study is to synthesize an efficacious system that can analyze images and distinguish between handwritten text and digits and then to translate the recognized text to several languages. For digit recognition, the CNNs are used while for text extraction, the OCR is employed, and the language translation capabilities are then embedded using Google Translate API. While the CNN model is trained on open datasets such as MNIST for examples of handwritten digits, datasets acquired from organizations are used for handwritten text. These steps include resizing, augmentation and normalization to avoid overfitting of the models. Hence, evaluation tools such as accuracy, precision, recall, and F-measure are used to measure the efficiency of the recognition and translation system. Some of the findings that have been found include the following: The results show that the proposed system has high accuracy rates in recognizing the handwritten digits and text where the average recognition accuracy was found to be above 95% for digits and 90% for text. The translation component enables the identified text to be translated into various languages apart from identifying them, thus expanding the utility of the system. This study enriches the knowledge about text recognition and opens possibilities for further studies in multiple languages handwritten text recognition and automated translation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhanced Handwritten Text and Digit Recognition with Multilingual Translation Using CNN and OCR Technologies

  • Nosirova Tanzila,
  • Rayyona Rustamjonova,
  • Khadichabonu Karimberdiyeva,
  • Gaurav Aggarwal,
  • Danish Ather,
  • Naina Chaudhary

摘要

The past few years have been marked by considerable growth of interest in handwritten recognition systems due to their usefulness in tasks such as automatic form processing, document scanning, and human–computer interface. The current paper offers an improved technique for text and handwriting digit identification through the incorporation of CNNs, OCR, and language translation technology. The aim of this study is to synthesize an efficacious system that can analyze images and distinguish between handwritten text and digits and then to translate the recognized text to several languages. For digit recognition, the CNNs are used while for text extraction, the OCR is employed, and the language translation capabilities are then embedded using Google Translate API. While the CNN model is trained on open datasets such as MNIST for examples of handwritten digits, datasets acquired from organizations are used for handwritten text. These steps include resizing, augmentation and normalization to avoid overfitting of the models. Hence, evaluation tools such as accuracy, precision, recall, and F-measure are used to measure the efficiency of the recognition and translation system. Some of the findings that have been found include the following: The results show that the proposed system has high accuracy rates in recognizing the handwritten digits and text where the average recognition accuracy was found to be above 95% for digits and 90% for text. The translation component enables the identified text to be translated into various languages apart from identifying them, thus expanding the utility of the system. This study enriches the knowledge about text recognition and opens possibilities for further studies in multiple languages handwritten text recognition and automated translation.