<p>Accurate RGB-D dense prediction requires jointly modeling global semantic consistency, local structural details, and reliable boundary cues. This becomes particularly challenging in dual-stream settings, where the RGB and depth images often differ in feature distributions and exhibit ambiguous boundaries. To address this issue, we propose a Dual-Domain Cross-Attention Fusion Network (DD-CAFNet) for enhanced feature fusion. Specifically, we introduce a Dual-Domain Confidence-Gated Cross Attention module (DD-CGCA), which adaptively fuses spectral and token-domain responses according to content-aware confidence-gated maps. The module consists of two cross-attention mechanisms: a spectral correlation branch that captures long-range structural correspondence in the frequency domain, and a token-aware spectral residual branch that performs semantic interaction in the spatial domain while preserving frequency-enhanced value representations. To further improve boundary recovery, we propose an Edge-Guided Frequency Decoupling module (EGFD) that decomposes features into low- and high-frequency components and dynamically routes them with edge priors. Finally, built upon a hierarchical dual-stream architecture, the proposed framework progressively aggregates cross-modal features from shallow to deep stages and enhances them using multi-scale feature refinement. Extensive experiments demonstrate that our model outperforms the state-of-the-art (SOTA) models qualitatively and quantitatively due to multi-domain context information. Our code is publicly available at: <a href="https://github.com/zhx-hub/DD-CAFNet.git.">https://github.com/zhx-hub/DD-CAFNet.git.</a></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dual-domain cross-attention fusion with edge-guided frequency decoupling for RGB-D saliency object detection

  • Xin Zhou,
  • Wenyao Ji,
  • Yongchao Zhang

摘要

Accurate RGB-D dense prediction requires jointly modeling global semantic consistency, local structural details, and reliable boundary cues. This becomes particularly challenging in dual-stream settings, where the RGB and depth images often differ in feature distributions and exhibit ambiguous boundaries. To address this issue, we propose a Dual-Domain Cross-Attention Fusion Network (DD-CAFNet) for enhanced feature fusion. Specifically, we introduce a Dual-Domain Confidence-Gated Cross Attention module (DD-CGCA), which adaptively fuses spectral and token-domain responses according to content-aware confidence-gated maps. The module consists of two cross-attention mechanisms: a spectral correlation branch that captures long-range structural correspondence in the frequency domain, and a token-aware spectral residual branch that performs semantic interaction in the spatial domain while preserving frequency-enhanced value representations. To further improve boundary recovery, we propose an Edge-Guided Frequency Decoupling module (EGFD) that decomposes features into low- and high-frequency components and dynamically routes them with edge priors. Finally, built upon a hierarchical dual-stream architecture, the proposed framework progressively aggregates cross-modal features from shallow to deep stages and enhances them using multi-scale feature refinement. Extensive experiments demonstrate that our model outperforms the state-of-the-art (SOTA) models qualitatively and quantitatively due to multi-domain context information. Our code is publicly available at: https://github.com/zhx-hub/DD-CAFNet.git.