Robust and Efficient Place Recognition via Camera-LiDAR Fusion and Semantic Descriptor
摘要
Fusion of multimodal information can improve the ability of robots and autonomous vehicles to perceive the environment. Descriptors generated from semantic objects and their relations have more robust performance. In this paper, we make the first attempt to generate semantic descriptors by fusing multimodal features, which are used for similarity comparison between scenes. We have designed two fusion models, one based on the camera view and one based on the bird’s eye view. Evaluation on the benchmark dataset shows that the fusion model based on the bird’s eye view exhibits improved recognition capability over prior methods. We first propose to use centroid and heatmap to generate semantic descriptors directly instead of clustering. The heatmap-based method drastically reduces the time from clustering, achieving generation in milliseconds and providing the possibility of real-time application.