Depth reliability-aware dual-stage network for RGB-D salient object segmentation
摘要
RGB-D salient object segmentation exploits the complementary characteristics of RGB appearance cues and depth geometric information, yet its performance is often hindered by noisy and spatially unreliable depth data as well as insufficient global-to-local refinement mechanisms. To address these issues, this paper proposes a depth reliability-aware dual-stage RGB-D salient object segmentation framework. The proposed method introduces a multi-level cross-modal feature fusion strategy that adaptively balances RGB semantic information and depth structural cues via modality-aware reweighting and residual semantic anchoring. On this basis, a coarse-to-fine segmentation paradigm is employed, where a lightweight decoder first generates a global coarse saliency mask that subsequently guides a mask-based refinement process to progressively enhance object boundaries and structural details under top-down semantic constraints. Extensive experiments on public RGB-D saliency benchmarks demonstrate that the proposed approach achieves strong structural consistency and competitive segmentation performance compared with recent state-of-the-art methods. In particular, the proposed framework shows advantages in preserving global object structure and maintaining stable pixel-level prediction under complex scenes with unreliable depth information.