Multimodal fusion and pseudo-labeling for enhanced weakly supervised camouflaged object detection
摘要
Camouflaged object detection (COD) is challenging because camouflaged objects are often highly similar to complex backgrounds. Existing COD methods usually rely on pixel-level annotations, which are labour-intensive and may involve subjective bias. To reduce the dependence on dense annotations, we propose SPMCNet, a weakly-supervised COD framework that incorporates generated pseudo-depth (PD) cues to complement RGB appearance information. Specifically, monocular depth estimation is integrated into the overall pipeline to generate PD maps from RGB images, which are used as auxiliary spatial cues rather than sensor-captured depth. To exploit RGB and PD information, we introduce a cross-modal adaptive relationship fusion (CARF) module, which models their relationship from spatial and channel views and adaptively adjusts their contributions during feature fusion. We further propose a cross-modal pseudo-supervised label mapping (CPLM) strategy to expand scribble annotations in both RGB and generated PD domains, generating pseudo-RGB labels and PD labels, respectively. These modality-specific pseudo-labels are then aggregated into pseudo-RGB-D (P-RGB-D) labels for training. In addition, a depth dual loss is introduced to supervise the auxiliary depth branch and encourage smoother object-region predictions. Experimental results on four benchmark datasets show that SPMCNet achieves competitive performance among weakly-supervised COD methods. Compared with representative weakly-supervised baselines, SPMCNet obtains average improvements of 3.2% in structure measure (