Arbitrary-Shape Text Spotting Based on Global, Pixel and Sequence Semantics
摘要
The field of end-to-end text spotting has garnered significant interest in recent years, propelled by the revealed intrinsic synergies between scene text detection and recognition. While advancements have been made, the challenge of arbitrarily shaped scene text spotting persists. This paper introduces an innovative feature augmentation module that addresses the issues of limited receptive fields and weak representation typical of lightweight backbone networks, while also enhancing multi-scale information more effectively and reducing information loss during feature aggregation. Furthermore, to extract a richer set of backbone features, we propose a dual information attention mechanism that enables the backbone network to neuronally focus on salient information.