<p>Natural Language Processing (NLP) plays a crucial role in understanding and processing text data, with applications across a wide range of languages. In this context, the task of handwritten Urdu text recognition presents significant challenges due to the intricacies of the script, including cursive connectivity, context-sensitive character forms, and variations in handwriting styles and stroke density. Existing methods struggle to capture fine-grained details and larger character spans, especially in low-resolution images. This paper introduces a Swin Transformer-based deep learning approach for handwritten Urdu text recognition, demonstrating significant improvements over traditional CNN models and Vision Transformer-based methods. The Swin Transformer utilizes window-based attention mechanisms to preserve essential features such as loops, dots (nuqta), and strokes, crucial for distinguishing visually similar characters. A multi-scale windowing technique ensures that both fine details and larger character structures are captured, enabling the model to handle overlapping characters and varying handwriting styles. By leveraging transformer-based architectures, the proposed model successfully adapts to the complexities of handwritten Urdu text. Experimental results on various handwritten Urdu datasets show that the proposed approach outperforms existing models in terms of recognition accuracy, lower error rates, and reduced computational costs. The model requires minimal storage and computational resources, with training times under three hours and memory usage below 12 GB, making it well-suited for resource-constrained environments. This research presents a novel approach to handwritten Urdu recognition and contributes the updated UHLD dataset for future research in the field.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Towards automatic efficient recognition of multiresolution offline handwritten Urdu text using swin transformers

  • Aejaz Farooq Ganai,
  • Nasir N. Hurrah,
  • Farida Khursheed

摘要

Natural Language Processing (NLP) plays a crucial role in understanding and processing text data, with applications across a wide range of languages. In this context, the task of handwritten Urdu text recognition presents significant challenges due to the intricacies of the script, including cursive connectivity, context-sensitive character forms, and variations in handwriting styles and stroke density. Existing methods struggle to capture fine-grained details and larger character spans, especially in low-resolution images. This paper introduces a Swin Transformer-based deep learning approach for handwritten Urdu text recognition, demonstrating significant improvements over traditional CNN models and Vision Transformer-based methods. The Swin Transformer utilizes window-based attention mechanisms to preserve essential features such as loops, dots (nuqta), and strokes, crucial for distinguishing visually similar characters. A multi-scale windowing technique ensures that both fine details and larger character structures are captured, enabling the model to handle overlapping characters and varying handwriting styles. By leveraging transformer-based architectures, the proposed model successfully adapts to the complexities of handwritten Urdu text. Experimental results on various handwritten Urdu datasets show that the proposed approach outperforms existing models in terms of recognition accuracy, lower error rates, and reduced computational costs. The model requires minimal storage and computational resources, with training times under three hours and memory usage below 12 GB, making it well-suited for resource-constrained environments. This research presents a novel approach to handwritten Urdu recognition and contributes the updated UHLD dataset for future research in the field.