In recent years, the field of autonomous vehicles and driverless technology has seen remarkable advancements, driven by contributions from mainstream automotive manufacturers and open-source projects. This research aims to develop a pipeline for road scene understanding through semantic segmentation. The proposed pipeline utilises a multi-modal segmentation model, incorporating greyscale images and point cloud data from Xenolidar, specifically designed to capture the structural priors of highway road scenes. The fusion of input modalities and the design of an encoder-decoder architecture with a novel attention scheme called HaWANet is introduced, which focuses on the height and width contextual information to improve the accuracy of road segmentation, are the primary aspects explored for the proposed model. The output of the encoder is a two-dimensional point cloud, which effectively represents the road’s planar nature, and is crucial for improving the accuracy of road segmentation, particularly in edge cases, addressing current challenges in autonomous driving research. This research, aimed at addressing the segmentation problem for multimodal sensor data, has presented significant performance improvement over single-modal approaches.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

HaWANet: Road Scene Understanding with Multi-modal Sensor Data Using Height-Width-Driven Attention Network

  • Soumick Chatterjee,
  • Jiahua Xu,
  • Adarsh Kuzhipathalil,
  • Andreas Nürnberger

摘要

In recent years, the field of autonomous vehicles and driverless technology has seen remarkable advancements, driven by contributions from mainstream automotive manufacturers and open-source projects. This research aims to develop a pipeline for road scene understanding through semantic segmentation. The proposed pipeline utilises a multi-modal segmentation model, incorporating greyscale images and point cloud data from Xenolidar, specifically designed to capture the structural priors of highway road scenes. The fusion of input modalities and the design of an encoder-decoder architecture with a novel attention scheme called HaWANet is introduced, which focuses on the height and width contextual information to improve the accuracy of road segmentation, are the primary aspects explored for the proposed model. The output of the encoder is a two-dimensional point cloud, which effectively represents the road’s planar nature, and is crucial for improving the accuracy of road segmentation, particularly in edge cases, addressing current challenges in autonomous driving research. This research, aimed at addressing the segmentation problem for multimodal sensor data, has presented significant performance improvement over single-modal approaches.