<p>Semantic segmentation has succeeded remarkably in various applications, such as autonomous vehicles and robotic systems. However, the training process for such techniques necessitates a significant amount of labeled data. Although semi-supervised frameworks can alleviate this issue, advanced approaches typically require multiple baseline models to form a dual model, which is costly in space computation complexity. To relieve the undesired computational cost for systems with precious computation and memory resources, we propose an efficient and scalable semi-supervised learning framework to significantly improve the performance with a few additional parameters concerning the baseline models. This framework includes a pseudo-dual module and a self-rectification module. The overall structure comprises three parts: an encoder, a shallow decoder, and a deep decoder. The deep decoder is connected to a deep layer of the encoder, and the shallow decoder is connected to a shallow layer. As knowledge distillation transfers knowledge from one model to another, the pseudo-dual module can distill knowledge from the ensemble of two decoders to improve the encoder, which can implicitly form a pseudo-dual model. The self-rectification module calculates class-wise likelihoods according to the similarity between features and class prototypes learned from different decoders and rectifies low-confidence pseudo-labels. The effectiveness of such rectification is justified theoretically and numerically. In our experiments with DeepLabV2, our methods outperform others in mIoU by over 1.21% with 1/8 labeled data using the Cityscapes dataset and by 0.38% with 1/8 labeled data using PASCAL VOC 2012 datasets. In most cases, our approach can also save more than 30% of memory costs during training. Nevertheless, the effectiveness of the proposed approach also depends on the quality of pseudo-labels generated by backbone models and may encounter challenges when handling data with highly imbalanced classes.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An efficient and scalable semi-supervised framework for semantic segmentation

  • Huazheng Hao,
  • Hui Xiao,
  • Junjie Xiong,
  • Li Dong,
  • Diqun Yan,
  • Dongtai Liang,
  • Jiayan Zhuang,
  • Chengbin Peng

摘要

Semantic segmentation has succeeded remarkably in various applications, such as autonomous vehicles and robotic systems. However, the training process for such techniques necessitates a significant amount of labeled data. Although semi-supervised frameworks can alleviate this issue, advanced approaches typically require multiple baseline models to form a dual model, which is costly in space computation complexity. To relieve the undesired computational cost for systems with precious computation and memory resources, we propose an efficient and scalable semi-supervised learning framework to significantly improve the performance with a few additional parameters concerning the baseline models. This framework includes a pseudo-dual module and a self-rectification module. The overall structure comprises three parts: an encoder, a shallow decoder, and a deep decoder. The deep decoder is connected to a deep layer of the encoder, and the shallow decoder is connected to a shallow layer. As knowledge distillation transfers knowledge from one model to another, the pseudo-dual module can distill knowledge from the ensemble of two decoders to improve the encoder, which can implicitly form a pseudo-dual model. The self-rectification module calculates class-wise likelihoods according to the similarity between features and class prototypes learned from different decoders and rectifies low-confidence pseudo-labels. The effectiveness of such rectification is justified theoretically and numerically. In our experiments with DeepLabV2, our methods outperform others in mIoU by over 1.21% with 1/8 labeled data using the Cityscapes dataset and by 0.38% with 1/8 labeled data using PASCAL VOC 2012 datasets. In most cases, our approach can also save more than 30% of memory costs during training. Nevertheless, the effectiveness of the proposed approach also depends on the quality of pseudo-labels generated by backbone models and may encounter challenges when handling data with highly imbalanced classes.