HaWANet: Road Scene Understanding with Multi-modal Sensor Data Using Height-Width-Driven Attention Network
摘要
In recent years, the field of autonomous vehicles and driverless technology has seen remarkable advancements, driven by contributions from mainstream automotive manufacturers and open-source projects. This research aims to develop a pipeline for road scene understanding through semantic segmentation. The proposed pipeline utilises a multi-modal segmentation model, incorporating greyscale images and point cloud data from Xenolidar, specifically designed to capture the structural priors of highway road scenes. The fusion of input modalities and the design of an encoder-decoder architecture with a novel attention scheme called HaWANet is introduced, which focuses on the height and width contextual information to improve the accuracy of road segmentation, are the primary aspects explored for the proposed model. The output of the encoder is a two-dimensional point cloud, which effectively represents the road’s planar nature, and is crucial for improving the accuracy of road segmentation, particularly in edge cases, addressing current challenges in autonomous driving research. This research, aimed at addressing the segmentation problem for multimodal sensor data, has presented significant performance improvement over single-modal approaches.