Manual annotation of undigitized manuscripts is a resource and labor intensive endeavour. On the other hand, using out-of-the-box Optical Character Recognition (OCR) tools to digitise text from handwritten manuscripts is challenging, particularly in languages like Sanskrit, where the scripts have evolved significantly over time. This poster presents an annotation tool which allows the user to extract text from undigitized manuscripts using OCR, following which users can make corrections to the OCR-detected text. Users can then request fine tuning on a few pages corrected by them, making the annotation process easier and more efficient for the subsequent pages by improving OCR performance. Additionally, an issue faced in the annotation of manuscripts is that the Devanāgari script is sometimes laborious to edit due to the behaviour of conjunct clusters and the halant, necessitating many keystrokes for simple edits. To mitigate this issue, the application uses an additional text box that represents the text in the Harvard-Kyoto transliteration scheme, which is easier to edit given that most keyboards use the Roman alphabet. The tool also supports rare Sanskrit characters, which are not usually supported by standard typing tools. The frontend, which is designed to reduce effort on the annotator’s part by placing the recognised text in editable text boxes right below the line images, is implemented in Vue.js. The backend is a Flask server employing PyTorch for inference and fine-tuning. The application aims to facilitate efficient annotation by having users upload manuscript images, which are then processed by the backend, which segments text lines from the leaves and recognises the text in the line images. Initial evaluations from beta testers suggest that the application significantly enhances the efficiency of the annotation process, making it more accessible and reducing the time required for its completion.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Semi-Automatic Text Recognition Tool for Pre-Colonial Handwritten Manuscripts in Devanāgari Script

  • Bharath Valaboju,
  • Shagun Dwivedi,
  • Kartik Chincholikar,
  • Kaushik Gopalan,
  • Vinod Vidwans

摘要

Manual annotation of undigitized manuscripts is a resource and labor intensive endeavour. On the other hand, using out-of-the-box Optical Character Recognition (OCR) tools to digitise text from handwritten manuscripts is challenging, particularly in languages like Sanskrit, where the scripts have evolved significantly over time. This poster presents an annotation tool which allows the user to extract text from undigitized manuscripts using OCR, following which users can make corrections to the OCR-detected text. Users can then request fine tuning on a few pages corrected by them, making the annotation process easier and more efficient for the subsequent pages by improving OCR performance. Additionally, an issue faced in the annotation of manuscripts is that the Devanāgari script is sometimes laborious to edit due to the behaviour of conjunct clusters and the halant, necessitating many keystrokes for simple edits. To mitigate this issue, the application uses an additional text box that represents the text in the Harvard-Kyoto transliteration scheme, which is easier to edit given that most keyboards use the Roman alphabet. The tool also supports rare Sanskrit characters, which are not usually supported by standard typing tools. The frontend, which is designed to reduce effort on the annotator’s part by placing the recognised text in editable text boxes right below the line images, is implemented in Vue.js. The backend is a Flask server employing PyTorch for inference and fine-tuning. The application aims to facilitate efficient annotation by having users upload manuscript images, which are then processed by the backend, which segments text lines from the leaves and recognises the text in the line images. Initial evaluations from beta testers suggest that the application significantly enhances the efficiency of the annotation process, making it more accessible and reducing the time required for its completion.