Scene text recognition is the task of identifying text in natural scene images. Popular scene text recognition technologies mostly employ Transformer-based encoder-decoder methods tailored for resource-rich languages. However, when training data is insufficient, Transformer-based methods perform poorly compared to CNN-based methods. Nonetheless, CNN-based methods cannot process global information and often fall short in capturing contextual information between characters. This paper proposes a Dual Feature Enhanced Scene Text Recognition Method based on the CNN encoder-decoder designed for low-resource Uyghur. In the encoder, we employ a dynamic attention enhancement technique to strengthen the model’s learning capability of features in both spatial and channel dimensions, reducing the model’s dependency on large-scale training data. In the decoder, we introduce a novel global feature enhancement strategy that associates features from the encoder globally, mitigating the convolutional neural network’s lack of global information processing ability. Additionally, we construct two Uyghur language scene datasets, named U1 and U2. Comparative experimental results demonstrate the outstanding performance of our method on the U1 and U2 datasets. Compared to baseline methods, our approach achieves a respective increase in accuracy of 5.2% and 3.2% while reducing model parameters.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dual Feature Enhanced Scene Text Recognition Method for Low-Resource Uyghur

  • Miaomiao Xu,
  • Jiang Zhang,
  • Lianghui Xu,
  • Yanbing Li,
  • Wushour Silamu

摘要

Scene text recognition is the task of identifying text in natural scene images. Popular scene text recognition technologies mostly employ Transformer-based encoder-decoder methods tailored for resource-rich languages. However, when training data is insufficient, Transformer-based methods perform poorly compared to CNN-based methods. Nonetheless, CNN-based methods cannot process global information and often fall short in capturing contextual information between characters. This paper proposes a Dual Feature Enhanced Scene Text Recognition Method based on the CNN encoder-decoder designed for low-resource Uyghur. In the encoder, we employ a dynamic attention enhancement technique to strengthen the model’s learning capability of features in both spatial and channel dimensions, reducing the model’s dependency on large-scale training data. In the decoder, we introduce a novel global feature enhancement strategy that associates features from the encoder globally, mitigating the convolutional neural network’s lack of global information processing ability. Additionally, we construct two Uyghur language scene datasets, named U1 and U2. Comparative experimental results demonstrate the outstanding performance of our method on the U1 and U2 datasets. Compared to baseline methods, our approach achieves a respective increase in accuracy of 5.2% and 3.2% while reducing model parameters.