Hybrid Encoding Method for Scene Text Recognition in Low-Resource Uyghur
摘要
Current advanced methods for scene text recognition are predominantly based on Transformer architecture, focusing primarily on resource-rich languages such as Chinese and English. However, Transformer-based architectures heavily rely on annotated data and their performance on low-resource data is not satisfactory. This paper proposes a Hybrid Encoding Method (HEM) for Scene Text Recognition in Low-Resource Uyghur, aiming to equip the network with both the long-range association capability of the Transformer for global image context and the function of CNN for capturing local detailed information. Simultaneously, by combining the strengths of CNN and Transformer encodings, the model’s learning capacity can be enhanced in low-resource settings, bolstering its ability to comprehend images while reducing its reliance on annotated data. On the other hand, we construct two Uyghur scene text datasets, namely U1 and U2. Experimental results demonstrate that the proposed hybrid encoding method achieves outstanding performance in low-resource Uyghur scene text recognition, improving accuracy by 15% compared to baseline methods.