Towards the digitization of Kurdish handwritten script using deep learning: a comprehensive analysis
摘要
Research on offline Central Kurdish (Sorani) handwritten text recognition remains substantially limited compared to other scripts, with most existing studies primarily focusing on isolated character or digit recognition rather than text-line or document-level analysis. The lack of a comprehensive review constitutes a significant factor contributing to this gap in Kurdish handwritten recognition research. This survey paper addresses this gap by critically evaluating the limitations of existing Kurdish handwritten datasets and models, encompassing all relevant works published to date. Furthermore, it systematically analyzes neural network approaches proposed in the literature for sequence-to-sequence handwritten recognition in languages with scripts similar to Kurdish, including Arabic, Persian, and Urdu. Common architectures employing modern approaches for handwritten text recognition, such as CNNs, RNNs, attention mechanisms, and Transformers, are described. The study highlights the distinctive features of the Kurdish script that pose unique recognition challenges, including contextual character variations, ligatures, and diacritical marks. A comparative analysis is provided between recognition methods for Kurdish and other Arabic-like scripts versus Latin-based handwritten recognition systems. The findings emphasize the need for further research in Kurdish and Arabic-like handwritten recognition and identify potential real-world applications enabled by such advancements, including processing bank checks, historical documents, government records, and court cases. Additionally, the discussion explores how future research might integrate generative models with text recognizers to enhance performance on limited training data and improve model generalization for page-level handwritten text processing.