Multi-stream information complementarity network for RGB-D camouflaged object detection
摘要
Camouflage object detection (COD) aims to identify objects hidden in natural scenes. Due to the imperceptible differences between disguised objects and their surrounding environment, accurately detecting disgusted objects is extremely challenging. To overcome this challenge, introducing depth map has become an important breakthrough in that it can provide valuable spatial clues for COD tasks. In this paper, we propose a multi-stream information complementarity network (MICNet) for camouflage object detection, comprehensively exploiting complementary useful information from depth and RGB images to boost the accuracy of detection. The MICNet adopts an encoding-fusion-decoding structure and utilizes a two-stream pyramid visual transformer as backbone to extract multi-level features from input images. In advance, coarse depth map is generated by a frozen depth estimation model for the given RGB image. We first propose a cross-modal attention enhancement module that adaptively integrates RGB features and depth features based on attention mechanisms. Then, we introduce a simple and effective positioning guidance generation module, which generates coarse positioning maps to provide localization guidance for subsequent decoding processes. Finally, we design a multi-stream information aggregation unit, which takes richer intermediate information into account to preserve spatial detail of targets. Extensive experimental results on four COD datasets indicate that our network achieves superior performance in four common metrics compared to 14 state-of-the-art COD methods, while also performing well in salient object detection tasks.