Transformer architectures have sparked significant interest in the field of end-to-end text detection and recognition, also known as text spotting. Given that most existing methods rely on CNN-based techniques, it opens up a considerable scope for exploration in transformer-based methods. Currently, transformer-based methods can be divided into two types: one uses two independent decoders for detection and recognition respectively, and the other incorporates the detection and recognition tasks into a single decoder. The first approach lacks interaction between detection and recognition, ignoring the synergy between these two tasks. The second approach, while providing an implicit interaction between two tasks in one decoder, overlooks their individual characteristics. In this paper, we consider the distinct attributes of detection and recognition as well as their integrated properties, and design a deep interaction module for bottom-up explicit hierarchical interaction between these tasks. Furthermore, to address the inconsistency in length between ground truth and text scripts, we introduce CTC loss in our method. This change provides a more efficient and accurate way to handle these inconsistencies. Although our method is conceptually simple, it outperforms the state-of-the-art methods on Total-Text, SCUT-CTW1500 and ICDAR15 benchmarks.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DeepTTS: Enhanced Transformer-Based Text Spotter via Deep Interaction Between Detection and Recognition Tasks

  • Yu Xie,
  • Canhui Xu,
  • Cao Shi,
  • Jialiang Li,
  • Zhengyi Yuan,
  • Qian Qiao

摘要

Transformer architectures have sparked significant interest in the field of end-to-end text detection and recognition, also known as text spotting. Given that most existing methods rely on CNN-based techniques, it opens up a considerable scope for exploration in transformer-based methods. Currently, transformer-based methods can be divided into two types: one uses two independent decoders for detection and recognition respectively, and the other incorporates the detection and recognition tasks into a single decoder. The first approach lacks interaction between detection and recognition, ignoring the synergy between these two tasks. The second approach, while providing an implicit interaction between two tasks in one decoder, overlooks their individual characteristics. In this paper, we consider the distinct attributes of detection and recognition as well as their integrated properties, and design a deep interaction module for bottom-up explicit hierarchical interaction between these tasks. Furthermore, to address the inconsistency in length between ground truth and text scripts, we introduce CTC loss in our method. This change provides a more efficient and accurate way to handle these inconsistencies. Although our method is conceptually simple, it outperforms the state-of-the-art methods on Total-Text, SCUT-CTW1500 and ICDAR15 benchmarks.