Cross-modality Attentional Fusing Network for RGB-D salient object detection
摘要
RGB-D salient object detection (SOD) has made significant progress and demonstrated remarkable results. However, most RGB-D SOD methods fail to consider the differences between modalities and the differences across hierarchical levels within modalities, and directly fuse them. These strategies may cause the loss or redundancy of information, which can lead to a degradation in model detection performance. To tackle these issues, we propose a Cross-modality Attentional Fusing Network(CAFNet), which exploits the complementarity of features across modalities and combine hierarchical features to achieve bimodal and multi-level feature fusion to enhance SOD performance. Firstly, a Bidirectional Feature Interaction module is designed to fully capture the complementary relationship between two modalities in channel and spatial dimension and realizes bidirectional feature interaction. Next, we design a Multi-Scale Feature Progressive Fusion module to expand the multi-scale context information of semantically rich but low-resolution deep features with limited receptive fields. In addition, to address the semantic gap in cross-level feature fusion, we propose a Hierarchical Feature Refinement module to alleviate the inconsistency of the direct fusion from hierarchical features. Comprehensive experiments on 5 benchmarks demonstrate that our CAFNet outperforms typical state-of-the-art methods.