SIE-DepthNet: Semantic-Guided Monocular Depth Estimation for Dynamic Environment
摘要
Depth estimation is a fundamental aspect of spatial computation in virtual-real fusion contexts, wherein self-supervised monocular depth estimation presents a persistent challenge within the fields of computer vision and augmented/mixed reality (AR/XR) applications. However, existing methods face difficulties with moving objects, occlusions, and motion blur, resulting in inaccurate estimates, particularly in dynamic regions. Previous solutions either omit challenging areas in the training process or employ pseudo-depth labels, which still cannot fully address the issues. In this paper, we introduce our semantically implicit and explicit depth estimation network, SIE-DepthNet, a novel approach that leverages semantic information to improve monocular depth estimation in dynamic scenes. We articulate a Depth Semantic Feature Fusion Network (DSFFNet) to provide implicit guidance, ensuring consistent depth distributions across categories. Additionally, we propose a semantic-guided ranking loss that optimizes depth accuracy by aligning estimated values with segmentation boundaries, using uncertainty-aware weighting to mitigate the impact of segmentation noise. Extensive experimental results demonstrate that SIE-DepthNet achieves accurate depth map predictions, significantly enhancing performance in dynamic and static scenarios.