<p>Text detection in natural scenes, as a prerequisite for text reading, has rapidly advanced across various industries. However, in Mongolian-Chinese bilingual text scenes, Mongolian text follows a top-to-bottom layout, while Chinese text is typically arranged from left to right. There is a lack of detection methods that effectively integrate both. To address this limitation, we propose a text detection method specifically designed for Mongolian-Chinese bilingual scenes in natural environments. To enable joint detection of Mongolian and Chinese text, our approach constructs a bold text outline map using classification and regression branches to extract keypoints and edges. In parallel, a refine text character map is generated by a component branch that predicts character-level bounding boxes. These two representations are organically integrated, allowing for bidirectional interaction between point-edge structures and character-box features. In response to the substantial differences in script structure between Mongolian and Chinese, we apply a boundary point offset strategy to refine candidate bounding boxes, ensuring complete text instance coverage for both languages. To address the issue of blurred text images, we introduce a ClearText module for image restoration, which enhances the visibility of fine textual details and facilitates more accurate detection. Experimental results on four public datasets, as well as a custom Mongolian-Chinese bilingual dataset, validate the effectiveness of our method. Specifically, our approach achieves an F-measure of 84.6% on the custom bilingual dataset. On the ICDAR2015, MSRA-TD500, CTW1500, and Total-Text datasets, the method attains F-measures of 91.2%, 90.2%, 87.0%, and 90.4%, respectively. To promote further research in Mongolian-Chinese bilingual scenarios, we release a manually annotated natural scene text dataset containing 1,502 images featuring both Mongolian and Chinese text. The dataset is publicly available at <a href="https://github.com/1264sw/MACB-database">https://github.com/1264sw/MACB-database</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A text clarification and deep relational reasoning method for Mongolian-Chinese bilingual arbitrary-shaped scene text detection

  • Yuefeng Liu,
  • Liuxu Ding,
  • Yunong Liu

摘要

Text detection in natural scenes, as a prerequisite for text reading, has rapidly advanced across various industries. However, in Mongolian-Chinese bilingual text scenes, Mongolian text follows a top-to-bottom layout, while Chinese text is typically arranged from left to right. There is a lack of detection methods that effectively integrate both. To address this limitation, we propose a text detection method specifically designed for Mongolian-Chinese bilingual scenes in natural environments. To enable joint detection of Mongolian and Chinese text, our approach constructs a bold text outline map using classification and regression branches to extract keypoints and edges. In parallel, a refine text character map is generated by a component branch that predicts character-level bounding boxes. These two representations are organically integrated, allowing for bidirectional interaction between point-edge structures and character-box features. In response to the substantial differences in script structure between Mongolian and Chinese, we apply a boundary point offset strategy to refine candidate bounding boxes, ensuring complete text instance coverage for both languages. To address the issue of blurred text images, we introduce a ClearText module for image restoration, which enhances the visibility of fine textual details and facilitates more accurate detection. Experimental results on four public datasets, as well as a custom Mongolian-Chinese bilingual dataset, validate the effectiveness of our method. Specifically, our approach achieves an F-measure of 84.6% on the custom bilingual dataset. On the ICDAR2015, MSRA-TD500, CTW1500, and Total-Text datasets, the method attains F-measures of 91.2%, 90.2%, 87.0%, and 90.4%, respectively. To promote further research in Mongolian-Chinese bilingual scenarios, we release a manually annotated natural scene text dataset containing 1,502 images featuring both Mongolian and Chinese text. The dataset is publicly available at https://github.com/1264sw/MACB-database.