Optical Character Recognition, or OCR, is a tool that could apprehend textual content in photographs or scanned documents and convert it into system-readable textual content. To recognize and extract information from snap shots, optical character popularity (OCR) software program uses exclusive codes and codes to research the form, sample and grouping of characters. In 2005, Hewlett-Packard (HP) created the open supply software program Tesseract OCR (Optical person reputation). Its reason is to extract text from images and convert it into gadget-readable text. Tesseract can manage a couple of fonts, sizes and languages, making it clean to apply for OCR. In this newsletter, the authors provide an in-intensity evaluate of the usage of the Tesseract OCR engine in Indian languages.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Review of Usage of Tesseract OCR Engine with Vernacular Indian Languages

  • Kartik Joshi,
  • Harshal Arolkar

摘要

Optical Character Recognition, or OCR, is a tool that could apprehend textual content in photographs or scanned documents and convert it into system-readable textual content. To recognize and extract information from snap shots, optical character popularity (OCR) software program uses exclusive codes and codes to research the form, sample and grouping of characters. In 2005, Hewlett-Packard (HP) created the open supply software program Tesseract OCR (Optical person reputation). Its reason is to extract text from images and convert it into gadget-readable text. Tesseract can manage a couple of fonts, sizes and languages, making it clean to apply for OCR. In this newsletter, the authors provide an in-intensity evaluate of the usage of the Tesseract OCR engine in Indian languages.