<p>RGB-D salient object detection (SOD) aims to detect salient objects utilizing RGB images and depth images. Although the fully supervised RGB-D SOD has achieved higher detection precision at present, training the models requires plenty of pixel-level labels, which restricts the popularization and application of fully supervised RGB-D SOD in real scenarios to some extent. In recent years, weakly supervised RGB-D SOD that does not require using plenty of pixel-level labels has attracted extensive attention from scholars. Among the multiple types of weak labels adopted by weakly supervised RGB-D SOD, scribble labels are becoming increasingly popular due to their simplicity and flexibility. However, the pixel sparsity of scribble labels leads to the lack of semantic information and structure information of salient objects, which in turn severely weakens the performance of weakly supervised RGB-D SOD models. To address this issue, we propose a semantic interaction integration and position-aware network, and design a teacher dual-student framework based on it. Specifically, we first design a channel-splitting semantic interaction integration module to fully mine the salient semantic information of salient objects by interacting the self-modality and inter-modality information in the channel dimension as well as integrating the differentiated semantic information in RGB features and depth features. Then, we design a position-aware modality fusion module to capture the richer structure information of salient objects by deeply fusing the multi-modal information on the basis of learning the position information in RGB features and depth features. Based on the two modules and high-resolution network (HRNet), we design a semantic interaction integration and position-aware network (SIIPANet). Using three SIIPANets, we design a teacher dual-student framework to improve the quality of the generated pseudo labels by using both scribble labels and a small number of pixel-level labels, which effectively addresses the issue of lack of semantic information and structure information of salient objects caused by the pixel sparsity of scribble labels. Extensive experimental results on multiple datasets indicate that our method achieves state-of-the-art performance in weakly supervised RGB-D SOD.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A semantic interaction integration and position-aware network and a teacher dual-student framework for weakly supervised RGB-D salient object detection

  • Haishun Du,
  • Zeyu Li,
  • Kangyi Qiao,
  • Guanghui Zhang,
  • Chengsi Liu

摘要

RGB-D salient object detection (SOD) aims to detect salient objects utilizing RGB images and depth images. Although the fully supervised RGB-D SOD has achieved higher detection precision at present, training the models requires plenty of pixel-level labels, which restricts the popularization and application of fully supervised RGB-D SOD in real scenarios to some extent. In recent years, weakly supervised RGB-D SOD that does not require using plenty of pixel-level labels has attracted extensive attention from scholars. Among the multiple types of weak labels adopted by weakly supervised RGB-D SOD, scribble labels are becoming increasingly popular due to their simplicity and flexibility. However, the pixel sparsity of scribble labels leads to the lack of semantic information and structure information of salient objects, which in turn severely weakens the performance of weakly supervised RGB-D SOD models. To address this issue, we propose a semantic interaction integration and position-aware network, and design a teacher dual-student framework based on it. Specifically, we first design a channel-splitting semantic interaction integration module to fully mine the salient semantic information of salient objects by interacting the self-modality and inter-modality information in the channel dimension as well as integrating the differentiated semantic information in RGB features and depth features. Then, we design a position-aware modality fusion module to capture the richer structure information of salient objects by deeply fusing the multi-modal information on the basis of learning the position information in RGB features and depth features. Based on the two modules and high-resolution network (HRNet), we design a semantic interaction integration and position-aware network (SIIPANet). Using three SIIPANets, we design a teacher dual-student framework to improve the quality of the generated pseudo labels by using both scribble labels and a small number of pixel-level labels, which effectively addresses the issue of lack of semantic information and structure information of salient objects caused by the pixel sparsity of scribble labels. Extensive experimental results on multiple datasets indicate that our method achieves state-of-the-art performance in weakly supervised RGB-D SOD.