Script identification of text in natural scene images is challenging due to complex backgrounds, arbitrary orientations, different-sized characters, varying fonts, and multiple styles. Most existing methods are not effective in the presence of the above challenges. This paper introduces a new approach based on the Xception architecture and employing the log-polar transformed original image as an additional input, enabling the extraction of cues that are invariant to rotation, scaling, but are sensitive to script. The rationale behind the proposed work is that the combination of global features with text style features makes a significant difference in discriminating between different scripts. To combine the features extracted by Xception from the input image and log the polar transform of the input image, the proposed method introduces a style-enhanced fusion block. In addition, to further improve the performance of script identification, the proposed approach uses a new receptive channel selective focal attention module. Comparative evaluation results on three benchmark datasets, namely CVSI 2015, SIW-13, and MLe2e show that the proposed method outperforms the state-of-the-art in terms of classification rate.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

XLSI: A New Xception and Log Polar Transform Based Approach for Scene Text Script Identification

  • Ayush Roy,
  • Shivakumara Palaiahnakote,
  • Umapada Pal,
  • Apostolos Antonacopoulos,
  • Michael Blumenstein

摘要

Script identification of text in natural scene images is challenging due to complex backgrounds, arbitrary orientations, different-sized characters, varying fonts, and multiple styles. Most existing methods are not effective in the presence of the above challenges. This paper introduces a new approach based on the Xception architecture and employing the log-polar transformed original image as an additional input, enabling the extraction of cues that are invariant to rotation, scaling, but are sensitive to script. The rationale behind the proposed work is that the combination of global features with text style features makes a significant difference in discriminating between different scripts. To combine the features extracted by Xception from the input image and log the polar transform of the input image, the proposed method introduces a style-enhanced fusion block. In addition, to further improve the performance of script identification, the proposed approach uses a new receptive channel selective focal attention module. Comparative evaluation results on three benchmark datasets, namely CVSI 2015, SIW-13, and MLe2e show that the proposed method outperforms the state-of-the-art in terms of classification rate.