Arbitrary shaped scene text detection is a very challenging task and has become a research hotspot in the field of computer vision. Existing methods mostly use CNN to extract multi-scale features, followed by segmentation-based approaches to obtain probability maps and convert them into text boxes. However, due to the limitation of the CNN receptive field and the lack of global semantic information, the text detection network cannot model the correlation between different text instances, which limits the further improvement of text detection performance. In this paper, we propose a text detection network with a semantic feature extractor (SFENet), which can robustly detect irregular text in scene images. Firstly, we use Vision Transformer as the semantic feature extraction module to extract global semantic information and capture relationships between different text instances. Secondly, through the fusion module, semantic information is injected into multi-scale features and adaptively fused, resulting in accurate probability maps and text boxes of text regions. Finally, we conduct experiments on three public datasets: ICDAR2015, Total-Text and MSRA-TD500, and compare with other algorithms. The results demonstrate that SFENet achieves good detection results for arbitrary shape scene text, and performs well on all three public datasets.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SFENet: Arbitrary Shapes Scene Text Detection with Semantic Feature Extractor

  • Hongwei Chen,
  • Mengxi Cheng,
  • Tianshun Cheng,
  • Yun Xiao

摘要

Arbitrary shaped scene text detection is a very challenging task and has become a research hotspot in the field of computer vision. Existing methods mostly use CNN to extract multi-scale features, followed by segmentation-based approaches to obtain probability maps and convert them into text boxes. However, due to the limitation of the CNN receptive field and the lack of global semantic information, the text detection network cannot model the correlation between different text instances, which limits the further improvement of text detection performance. In this paper, we propose a text detection network with a semantic feature extractor (SFENet), which can robustly detect irregular text in scene images. Firstly, we use Vision Transformer as the semantic feature extraction module to extract global semantic information and capture relationships between different text instances. Secondly, through the fusion module, semantic information is injected into multi-scale features and adaptively fused, resulting in accurate probability maps and text boxes of text regions. Finally, we conduct experiments on three public datasets: ICDAR2015, Total-Text and MSRA-TD500, and compare with other algorithms. The results demonstrate that SFENet achieves good detection results for arbitrary shape scene text, and performs well on all three public datasets.