<p>Detecting text in wild is a challenging task since natural scene text is arbitrary-oriented, multi-scaled, and multilingual, meanwhile scene texts are always of complex background and variant aspect ratios as well. Existing methods usually use rectangular box annotation to detect text. The main problem of these methods are the performance degradation when identifying text with large changes across scales and words with arbitrary rotation angles. This paper presents the SEMFNet (Scene Text Detection Multi-Path Dynamic Fusion Network), a MultiPath Dynamic Fusion Network combined with a Scale Estimation Module, designed specifically for detecting text in natural scene imagery. First, a feature extraction module enhanced by a deformable attention mechanism is used to optimize feature representation and enhance the model’s ability to detect text features. Secondly, the scale estimation module is used to evaluate the importance of receptive field feature maps of different sizes, and the weights are adjusted to enhance the extraction ability of multi-scale features. Finally, the model combines different levels of feature representation to accurately locate the text position. The proposed method can detect multi-oriented, multiscaled, and multi-lingual scene text in both high accuracy and efficiency. We validate the proposed SEMFNet several standard popular benchmarks, the results of the experiments indicate that our approach attains performance at the forefront of current technology. Furthermore, we apply SEMFNet into multi-lingual scene text recognition task by combining it with a text recognizer CRNN, which also achieves a good performance.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SEMFNet: Multi-path dynamic fusion text detection model combined with scale estimation module

  • Junru Liao,
  • Junkun Li,
  • Jian Ji

摘要

Detecting text in wild is a challenging task since natural scene text is arbitrary-oriented, multi-scaled, and multilingual, meanwhile scene texts are always of complex background and variant aspect ratios as well. Existing methods usually use rectangular box annotation to detect text. The main problem of these methods are the performance degradation when identifying text with large changes across scales and words with arbitrary rotation angles. This paper presents the SEMFNet (Scene Text Detection Multi-Path Dynamic Fusion Network), a MultiPath Dynamic Fusion Network combined with a Scale Estimation Module, designed specifically for detecting text in natural scene imagery. First, a feature extraction module enhanced by a deformable attention mechanism is used to optimize feature representation and enhance the model’s ability to detect text features. Secondly, the scale estimation module is used to evaluate the importance of receptive field feature maps of different sizes, and the weights are adjusted to enhance the extraction ability of multi-scale features. Finally, the model combines different levels of feature representation to accurately locate the text position. The proposed method can detect multi-oriented, multiscaled, and multi-lingual scene text in both high accuracy and efficiency. We validate the proposed SEMFNet several standard popular benchmarks, the results of the experiments indicate that our approach attains performance at the forefront of current technology. Furthermore, we apply SEMFNet into multi-lingual scene text recognition task by combining it with a text recognizer CRNN, which also achieves a good performance.