Text-line segmentation is still considered challenging for complex background scene images. The success of text detection and recognition depends on the success of the text segmentation. This study presents a new method for text segmentation to facilitate reliable detection and recognition. Therefore, we introduce a new model called Pixel Correlation and Gaussian Attention Driven Network (PCGAUNet) for text segmentation. To extract pixel correlation, we modified the MultiResUnet architecture, which leverages pixel-wise correlation to effectively highlight foreground pixels. In addition, the proposed model utilizes the prior spatial statistics of bottleneck features to create a learnable Gaussian distribution, which guides the decoder for accurate text segmentation. Experimental results on three standard scene text segmentation datasets, ICDAR13 FST, Total Text, and COCO-TS, show that the proposed model outperforms existing methods. Furthermore, the results for the underwater dataset UTS-55 show that our model is robust and generic.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

PCGAUNet: Pixel Correlation and Gaussian Attention Driven Network for Text Segmentation

  • Ayush Roy,
  • Shivakumara Palaiahnakote,
  • Umapada Pal,
  • Apostolos Antonacopoulos,
  • Raghavendra Ramachandra

摘要

Text-line segmentation is still considered challenging for complex background scene images. The success of text detection and recognition depends on the success of the text segmentation. This study presents a new method for text segmentation to facilitate reliable detection and recognition. Therefore, we introduce a new model called Pixel Correlation and Gaussian Attention Driven Network (PCGAUNet) for text segmentation. To extract pixel correlation, we modified the MultiResUnet architecture, which leverages pixel-wise correlation to effectively highlight foreground pixels. In addition, the proposed model utilizes the prior spatial statistics of bottleneck features to create a learnable Gaussian distribution, which guides the decoder for accurate text segmentation. Experimental results on three standard scene text segmentation datasets, ICDAR13 FST, Total Text, and COCO-TS, show that the proposed model outperforms existing methods. Furthermore, the results for the underwater dataset UTS-55 show that our model is robust and generic.