<p>The growing demand for efficient document processing in Indic languages necessitates robust methods for analyzing document layouts, particularly for complex scripts like Devanagari. Need for digitization of Devanagari documents in fields such as education, historical research, and government operations has increased the demand for efficient and accurate layout analysis tools. This paper presents a novel approach utilizing a Devanagari Character Encoded Mix-Merge Vision Transformer (MViT) for comprehensive document layout analysis. Our proposed framework combines character encoding tailored for Devanagari with a unique mix-merge mechanism to capture spatial relationships and contextual information in document layouts effectively. The proposed model addresses the challenges of overlapping characters, Shirorekha, and varying text structures through advanced preprocessing and segmentation techniques. The approach leverages Devanagari OCR to segment characters, forming patches that are passed through embeddings to generate character encodings. These encodings are processed using a mix-merge mechanism, which captures spatial relationships and contextual information to reconstruct characters into sentences and paragraphs. Experimental results demonstrate the superior performance of the MViT model, achieving an average segmentation accuracy of 97.4%, and layout classification precision of 98.8%. The proposed framework demonstrates substantial improvements in the accuracies of segmentation and layout as compared with benchmark models, making it well-suited for Devanagari document digitization and analysis.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Devanagari character encoded mix-merge vision transformer for robust document layout analysis

  • Shweta Singh,
  • Sudeep Varshney,
  • Ankur Choudhary

摘要

The growing demand for efficient document processing in Indic languages necessitates robust methods for analyzing document layouts, particularly for complex scripts like Devanagari. Need for digitization of Devanagari documents in fields such as education, historical research, and government operations has increased the demand for efficient and accurate layout analysis tools. This paper presents a novel approach utilizing a Devanagari Character Encoded Mix-Merge Vision Transformer (MViT) for comprehensive document layout analysis. Our proposed framework combines character encoding tailored for Devanagari with a unique mix-merge mechanism to capture spatial relationships and contextual information in document layouts effectively. The proposed model addresses the challenges of overlapping characters, Shirorekha, and varying text structures through advanced preprocessing and segmentation techniques. The approach leverages Devanagari OCR to segment characters, forming patches that are passed through embeddings to generate character encodings. These encodings are processed using a mix-merge mechanism, which captures spatial relationships and contextual information to reconstruct characters into sentences and paragraphs. Experimental results demonstrate the superior performance of the MViT model, achieving an average segmentation accuracy of 97.4%, and layout classification precision of 98.8%. The proposed framework demonstrates substantial improvements in the accuracies of segmentation and layout as compared with benchmark models, making it well-suited for Devanagari document digitization and analysis.