<p>Automatic Train Operation (ATO) systems depend on multimodal sensors, such as camera and LiDAR, for reliable all-weather perception. However, conventional approaches that process modalities separately before fusion struggle with the heterogeneity of multimodal data, leading to suboptimal results. To address this, we propose a novel 3D semantic segmentation framework that integrates heterogeneous multimodal alignment with cross modal knowledge distillation, enabling effective collaboration between point cloud and image modalities. Our method leverages geometric priors for feature alignment, performs multi-stage cross-modal fusion for comprehensive scene understanding, and employs consistent collaborative distillation to enhance multimodal representation while preserving unimodal discriminative ability. Additionally, we construct a dedicated railway 3D multimodal dataset comprising approximately 75 paired frames across 10&#xa0;km of scenes. Experiments demonstrate that our approach surpasses state-of-the-art methods on the railway dataset and exhibits strong generalization to NuScenes and SemanticKITTI benchmarks.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

3D semantic segmentation for railway scenes via heterogeneous multimodal alignment and distillation

  • Ning Sun,
  • Yuchen Sun,
  • Maomao Sun,
  • Jixin Liu,
  • Lei Chai,
  • Cong Wu

摘要

Automatic Train Operation (ATO) systems depend on multimodal sensors, such as camera and LiDAR, for reliable all-weather perception. However, conventional approaches that process modalities separately before fusion struggle with the heterogeneity of multimodal data, leading to suboptimal results. To address this, we propose a novel 3D semantic segmentation framework that integrates heterogeneous multimodal alignment with cross modal knowledge distillation, enabling effective collaboration between point cloud and image modalities. Our method leverages geometric priors for feature alignment, performs multi-stage cross-modal fusion for comprehensive scene understanding, and employs consistent collaborative distillation to enhance multimodal representation while preserving unimodal discriminative ability. Additionally, we construct a dedicated railway 3D multimodal dataset comprising approximately 75 paired frames across 10 km of scenes. Experiments demonstrate that our approach surpasses state-of-the-art methods on the railway dataset and exhibits strong generalization to NuScenes and SemanticKITTI benchmarks.