Recently, Referring Image Segmentation (RIS) has gained significant attention for segmenting objects in images based on natural language descriptions. This task integrates computer vision and natural language understanding. Current methods mostly use Transformer as a visual encoder, which is good at modeling long-distance dependencies but easily ignores local details; while CNN can extract rich local features but struggles with capturing global information. For this reason, this paper proposes CNNFormer, a hybrid CNN-Transformer architecture, to fully leverage the complementary advantages of the two to enhance segmentation performance. In the encoding stage, CNN and Transformer extract multimodal features separately and optimize visual and language feature interactions with a well-designed multi-scale cross-modal attention module. Additionally, this paper presents a hierarchical feature fusion module that enhanc-es the integration of local and global information, aiming to improve segmentation performance. The experimental results show that CNNFormer improves 3.91% over a single CNN model and 2.97% over a single Transformer model. The effectiveness of the CNNFormer architecture in RIS tasks is verified and provides new research directions in this field.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CNNFormer: A CNN-Transformer Hybrid Model for Referring Image Segmentation

  • Kangsai Yao,
  • Guang Feng,
  • Xizhan Gao,
  • Xiaofeng Qu,
  • Sijie Niu

摘要

Recently, Referring Image Segmentation (RIS) has gained significant attention for segmenting objects in images based on natural language descriptions. This task integrates computer vision and natural language understanding. Current methods mostly use Transformer as a visual encoder, which is good at modeling long-distance dependencies but easily ignores local details; while CNN can extract rich local features but struggles with capturing global information. For this reason, this paper proposes CNNFormer, a hybrid CNN-Transformer architecture, to fully leverage the complementary advantages of the two to enhance segmentation performance. In the encoding stage, CNN and Transformer extract multimodal features separately and optimize visual and language feature interactions with a well-designed multi-scale cross-modal attention module. Additionally, this paper presents a hierarchical feature fusion module that enhanc-es the integration of local and global information, aiming to improve segmentation performance. The experimental results show that CNNFormer improves 3.91% over a single CNN model and 2.97% over a single Transformer model. The effectiveness of the CNNFormer architecture in RIS tasks is verified and provides new research directions in this field.