<p>To enable outdoor mobile robots to understand their surroundings, RGB-D semantic segmentation (SS)—which leverages both color (RGB) and depth (D) information—is an effective approach. Real-time performance is essential in this context, yet most existing RGB-D SS methods struggle with high-resolution inputs due to complex architectures and costly depth preprocessing. We propose a real-time RGB-D SS network by extending an existing real-time RGB SS model with depth integration, achieving both speed and accuracy. Our method introduces three key modules: 1) a scale-invariant depth encoder (SIDE) for efficient depth feature extraction, 2) an attentive feature fusion module (AFFM) for attention-based RGB-D fusion, and 3) a noise robust guiding module (NRGM) to handle depth noise without extra inference cost. By applying these modules to RGB networks, we create RGB-D models that match state-of-the-art accuracy while running at nearly one-third the computation time.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Real-time RGB-D Semantic Segmentation With Scale-invariant Depth Encoding and Noise-robust Fusion

  • Suhan Woo,
  • Junhyuk Hyun,
  • Suhyeon Lee,
  • Euntai Kim

摘要

To enable outdoor mobile robots to understand their surroundings, RGB-D semantic segmentation (SS)—which leverages both color (RGB) and depth (D) information—is an effective approach. Real-time performance is essential in this context, yet most existing RGB-D SS methods struggle with high-resolution inputs due to complex architectures and costly depth preprocessing. We propose a real-time RGB-D SS network by extending an existing real-time RGB SS model with depth integration, achieving both speed and accuracy. Our method introduces three key modules: 1) a scale-invariant depth encoder (SIDE) for efficient depth feature extraction, 2) an attentive feature fusion module (AFFM) for attention-based RGB-D fusion, and 3) a noise robust guiding module (NRGM) to handle depth noise without extra inference cost. By applying these modules to RGB networks, we create RGB-D models that match state-of-the-art accuracy while running at nearly one-third the computation time.