Few-shot semantic segmentation (FSS) aims to perform pixel-level segmentation on new object categories with limited labeled samples. Existing prototype-based methods struggle with low-confidence regions, blurry boundaries, and similar foreground and background, leading to poor prototype matching and reduced segmentation performance. To address this, we propose a dual-module framework consisting of the multi-modal support prototype enhancement (MSPE) and confusion region mining (CRM) modules. The MSPE module integrates query features into support features across channel and spatial dimensions and combines text embeddings from CLIP to build more expressive foreground and background prototypes. The CRM module enhances the query prototype by using multi-modal support prototypes to identify ambiguous regions, constructing auxiliary prototypes through similarity redistribution to improve the query prototype’s semantic representation and enhance the model’s ability to perceive complex regions. Experimental results show that our method achieves excellent performance on the PASCAL- \({5}^{i}\) and COCO- \({20}^{\text{i}}\) datasets.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Research on Multi-modal Prototype Guided Confusion Region Mining Method for Few-Shot Segmentation

  • Ruizhe Zhang,
  • Yanjun Yin,
  • Min Zhi,
  • Qiaozhi Xu,
  • Shiyuan Liu,
  • Yuxin Xia

摘要

Few-shot semantic segmentation (FSS) aims to perform pixel-level segmentation on new object categories with limited labeled samples. Existing prototype-based methods struggle with low-confidence regions, blurry boundaries, and similar foreground and background, leading to poor prototype matching and reduced segmentation performance. To address this, we propose a dual-module framework consisting of the multi-modal support prototype enhancement (MSPE) and confusion region mining (CRM) modules. The MSPE module integrates query features into support features across channel and spatial dimensions and combines text embeddings from CLIP to build more expressive foreground and background prototypes. The CRM module enhances the query prototype by using multi-modal support prototypes to identify ambiguous regions, constructing auxiliary prototypes through similarity redistribution to improve the query prototype’s semantic representation and enhance the model’s ability to perceive complex regions. Experimental results show that our method achieves excellent performance on the PASCAL- \({5}^{i}\) and COCO- \({20}^{\text{i}}\) datasets.