A Review of Usage of Tesseract OCR Engine with Vernacular Indian Languages
摘要
Optical Character Recognition, or OCR, is a tool that could apprehend textual content in photographs or scanned documents and convert it into system-readable textual content. To recognize and extract information from snap shots, optical character popularity (OCR) software program uses exclusive codes and codes to research the form, sample and grouping of characters. In 2005, Hewlett-Packard (HP) created the open supply software program Tesseract OCR (Optical person reputation). Its reason is to extract text from images and convert it into gadget-readable text. Tesseract can manage a couple of fonts, sizes and languages, making it clean to apply for OCR. In this newsletter, the authors provide an in-intensity evaluate of the usage of the Tesseract OCR engine in Indian languages.