Weakly Supervised Semantic Segmentation (WSSS) methods based on image-level labeling alleviate the burden of annotation due to image-level labels only providing object categories. However, the generated Class Activation Maps (CAM) can only localize the most discriminative region rather than the entire object region. To remedy this issue, a framework of WSSS algorithms for Cross-Domain Calibration and Boundary Denoising Network (CD-CBN) is presented. Specifically, a Spatial Feature Calibration Network (SFCN) is proposed to align cross-dimensional features with class prototypes, focusing on intra-class feature consistency. Then, a Class-Specific Distance Model (CSDM) is adopted to separate features from different classes, and feature activation in the object region surpasses the background area. Finally, a Full-domain-aware Noise Reduction Model (FNRM) is designed to refine the object boundary pixels by capturing global contextual features and further filtering out pixel-level noise. A comprehensive experimental evaluation of the highly challenging Pascal VOC 2012 dataset and MS COCO 2014 is presented in this study, illustrating the effectiveness of our suggested approach.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Cross-Domain Calibration and Boundary Denoising Network for Weakly Supervised Semantic Segmentation

  • Zhoufeng Liu,
  • Bingrui Li,
  • Shumin Ding,
  • Jiangtao Xi,
  • Chunlei Li

摘要

Weakly Supervised Semantic Segmentation (WSSS) methods based on image-level labeling alleviate the burden of annotation due to image-level labels only providing object categories. However, the generated Class Activation Maps (CAM) can only localize the most discriminative region rather than the entire object region. To remedy this issue, a framework of WSSS algorithms for Cross-Domain Calibration and Boundary Denoising Network (CD-CBN) is presented. Specifically, a Spatial Feature Calibration Network (SFCN) is proposed to align cross-dimensional features with class prototypes, focusing on intra-class feature consistency. Then, a Class-Specific Distance Model (CSDM) is adopted to separate features from different classes, and feature activation in the object region surpasses the background area. Finally, a Full-domain-aware Noise Reduction Model (FNRM) is designed to refine the object boundary pixels by capturing global contextual features and further filtering out pixel-level noise. A comprehensive experimental evaluation of the highly challenging Pascal VOC 2012 dataset and MS COCO 2014 is presented in this study, illustrating the effectiveness of our suggested approach.