To address the significant appearance differences between support and query samples from the same class in few-shot semantic segmentation, we introduce a feature-guided prototype augmentation model. This model incorporates a dual-branch feature enhancement module that combines the local sensitivity of cross-attention mechanisms with the global perception ability of Gram matrices, significantly enhancing the information exchange between the support set and the query set, leading to precise alignment of both local and global features. In addition, we propose a Hybrid Background Attention (HBA) module designed to more effectively model background characteristics. This module leverages both spatial and channel-wise information to refine background representations. By capturing the relative positional relationships among pixels, it improves the local semantic coherence of background prototypes. Furthermore, it dynamically assesses the significance of each feature channel, allowing for a more precise encoding of background semantics. We evaluate the proposed approach on two widely used benchmarks: PASCAL- \({5}^{i}\) and COCO- \({20}^{\text{i}}\) . Experimental results demonstrate that the proposed method exhibits exceptional meta-learning capabilities, significantly improving the performance of few-shot semantic segmentation tasks.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Feature-Guided Prototype-Enhanced Few-Shot Semantic Segmentation Model

  • Shiyuan Liu,
  • Yanjun Yin,
  • Min Zhi,
  • Qiaozhi Xu,
  • Ruizhe Zhang,
  • Yuxin Xia

摘要

To address the significant appearance differences between support and query samples from the same class in few-shot semantic segmentation, we introduce a feature-guided prototype augmentation model. This model incorporates a dual-branch feature enhancement module that combines the local sensitivity of cross-attention mechanisms with the global perception ability of Gram matrices, significantly enhancing the information exchange between the support set and the query set, leading to precise alignment of both local and global features. In addition, we propose a Hybrid Background Attention (HBA) module designed to more effectively model background characteristics. This module leverages both spatial and channel-wise information to refine background representations. By capturing the relative positional relationships among pixels, it improves the local semantic coherence of background prototypes. Furthermore, it dynamically assesses the significance of each feature channel, allowing for a more precise encoding of background semantics. We evaluate the proposed approach on two widely used benchmarks: PASCAL- \({5}^{i}\) and COCO- \({20}^{\text{i}}\) . Experimental results demonstrate that the proposed method exhibits exceptional meta-learning capabilities, significantly improving the performance of few-shot semantic segmentation tasks.