A Coordinated Interaction and Cross Enhancement Network for Salient Object Detection
摘要
The effective fusion of information from RGB modality and depth modality is currently a prominent research area within the field of RGB-D salient object detection. However, most of the existing RGB-D salient object detection methods not only fail to fully interact RGB features with depth features before fusing them, but also fail to utilize the connection between RGB modality and depth modality to enhance the features, which limits the performance of their models. To address the above problems, we propose a coordinated interaction and cross enhancement network for RGB-D salient object detection. Specifically, we first design a feature coordinated interaction module, which facilitates coordinated interaction between RGB features and depth features by manipulating information in both the channel and spatial dimensions to mitigate the adverse impact of modality differences. Next, we design a cross-modality cross-enhancement module to enhance the salient information in RGB features and depth features by exploiting the connection between the RGB modality and the depth modality based on capturing the long range contextual information between the two modalities. Furthermore, we design a two-stage feature fusion module to generate fused features containing rich saliency information by fusing RGB features and depth features. Finally, we decode the fused features step by step to get the final saliency map. Comprehensive experimental results on five public datasets verify that our model surpasses the performance of a majority of state of the art RGB-D salient object detection models.