<p>Multi-modal 3D object detection methods combine data from multiple sensors, like cameras and LiDAR, to improve environmental perception. However, these methods are confined to learning representations within a closed vocabulary context and struggle with the exploration required for the open-vocabulary scenarios, thus hindering their broad application. In this paper, we propose a novel method SF-CLIP3D by leveraging spatial-frequency-enhanced vision-language models (e.g., CLIP) to address the above limitation for multi-modal 3D object detection. Concretely, we leverage CLIP’s open-vocabulary capabilities and utilize it as the backbone network, and then we refine it with additional multi-scale and detailed geometric information, thereby significantly improving the model’s generalization and bridging the gap between 2 and 3D tasks. Moreover, we propose an innovative bilateral-frequency module that integrates high-level spatial and frequency-domain information to bolster the robustness of image features, thus diminishing noise and enriching the feature representations for multi-modal 3D object detection. We also employ a cross-modality fusion strategy to effectively integrate image and point cloud features. Extensive experiments on the KITTI dataset demonstrate the effectiveness of our proposed method. Notably, our method achieves a 2.96% mAP increase over the second-best competitor on the testing dataset.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SF-CLIP3D: spatial-frequency-enhanced vision-language models for multi-modal 3D object detection

  • Haoxin Gu,
  • Ruiping Xiong,
  • Yu Tian

摘要

Multi-modal 3D object detection methods combine data from multiple sensors, like cameras and LiDAR, to improve environmental perception. However, these methods are confined to learning representations within a closed vocabulary context and struggle with the exploration required for the open-vocabulary scenarios, thus hindering their broad application. In this paper, we propose a novel method SF-CLIP3D by leveraging spatial-frequency-enhanced vision-language models (e.g., CLIP) to address the above limitation for multi-modal 3D object detection. Concretely, we leverage CLIP’s open-vocabulary capabilities and utilize it as the backbone network, and then we refine it with additional multi-scale and detailed geometric information, thereby significantly improving the model’s generalization and bridging the gap between 2 and 3D tasks. Moreover, we propose an innovative bilateral-frequency module that integrates high-level spatial and frequency-domain information to bolster the robustness of image features, thus diminishing noise and enriching the feature representations for multi-modal 3D object detection. We also employ a cross-modality fusion strategy to effectively integrate image and point cloud features. Extensive experiments on the KITTI dataset demonstrate the effectiveness of our proposed method. Notably, our method achieves a 2.96% mAP increase over the second-best competitor on the testing dataset.