In this work, the effectiveness of the open-source Tesseract OCR engine in identifying Dari scripts from a wide variety of photos is investigated, including academic transcripts, license plates, advertisements, non-linear text, etc. Because Dari is written in modified Arabic script, its cursive style, contextual letter forms, and lack of digital resources make it particularly challenging for OCR tasks. In order to improve text visibility, images were taken using a smartphone and processed using a specially designed preprocessing pipeline that included de-skewing, binarization, noise reduction, and grayscale conversion. To extract text from the photos, we set Tesseract using the Farsi (fas) language model. The results were assessed using error rates for words and characters, with post-processing used to correct typical recognition errors. The findings show that although Tesseract can extract Dari text with a decent level of accuracy, script variation and picture quality have a significant impact on performance. The paper offers suggestions for future research and illustrates the advantages and limitations of open-source OCR Tesseract for low-resource languages.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Optical Character Recognition for Dari Language in Afghanistan

  • Sadullah Karimi,
  • Annajiat Alim Rasel

摘要

In this work, the effectiveness of the open-source Tesseract OCR engine in identifying Dari scripts from a wide variety of photos is investigated, including academic transcripts, license plates, advertisements, non-linear text, etc. Because Dari is written in modified Arabic script, its cursive style, contextual letter forms, and lack of digital resources make it particularly challenging for OCR tasks. In order to improve text visibility, images were taken using a smartphone and processed using a specially designed preprocessing pipeline that included de-skewing, binarization, noise reduction, and grayscale conversion. To extract text from the photos, we set Tesseract using the Farsi (fas) language model. The results were assessed using error rates for words and characters, with post-processing used to correct typical recognition errors. The findings show that although Tesseract can extract Dari text with a decent level of accuracy, script variation and picture quality have a significant impact on performance. The paper offers suggestions for future research and illustrates the advantages and limitations of open-source OCR Tesseract for low-resource languages.