3D semantic segmentation for railway scenes via heterogeneous multimodal alignment and distillation
摘要
Automatic Train Operation (ATO) systems depend on multimodal sensors, such as camera and LiDAR, for reliable all-weather perception. However, conventional approaches that process modalities separately before fusion struggle with the heterogeneity of multimodal data, leading to suboptimal results. To address this, we propose a novel 3D semantic segmentation framework that integrates heterogeneous multimodal alignment with cross modal knowledge distillation, enabling effective collaboration between point cloud and image modalities. Our method leverages geometric priors for feature alignment, performs multi-stage cross-modal fusion for comprehensive scene understanding, and employs consistent collaborative distillation to enhance multimodal representation while preserving unimodal discriminative ability. Additionally, we construct a dedicated railway 3D multimodal dataset comprising approximately 75 paired frames across 10 km of scenes. Experiments demonstrate that our approach surpasses state-of-the-art methods on the railway dataset and exhibits strong generalization to NuScenes and SemanticKITTI benchmarks.