In recent years, cooperative perception (CP) in vehicle-to-infrastructure (V2I) scenarios has gained significant traction as a key technology in autonomous driving. In this paper, we investigate the end-to-end object detection model and spatiotemporal asynchrony to enhance the perception performance of autonomous vehicles. We propose a novel V2I CP framework termed V2ICooper, designed for efficient and robust object detection and fusion. We propose an end-to-end object detection model with a heterogeneous multi-agent middle layer (HMML) serving as a backbone module. HMML facilitates feature interaction across different levels, allowing for the exploration of richer features and enhancing the system’s detection performance. To mitigate the impact of spatiotemporal asynchrony on the results, we introduce the spatiotemporal asynchronous fusion (SAF) method. This approach involves learning complex nonlinear mapping relationships between input sequences and corresponding object sequences, enabling spatiotemporal alignment. Experimental validations conducted by V2ICooper on real-world DAIR-V2X-C dataset demonstrate superior accuracy and robustness in object detection. Additionally, the successful implementation of the proposed system in real scenarios substantiates its effectiveness, as evidenced by experimental results.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

V2ICooper: Toward Vehicle-to-Infrastructure Cooperative Perception with Spatiotemporal Asynchronous Fusion

  • Sheng Yi,
  • Hao Zhang,
  • Feiyu Jin,
  • Yiyang Hu,
  • Rongzhen Li,
  • Kai Liu

摘要

In recent years, cooperative perception (CP) in vehicle-to-infrastructure (V2I) scenarios has gained significant traction as a key technology in autonomous driving. In this paper, we investigate the end-to-end object detection model and spatiotemporal asynchrony to enhance the perception performance of autonomous vehicles. We propose a novel V2I CP framework termed V2ICooper, designed for efficient and robust object detection and fusion. We propose an end-to-end object detection model with a heterogeneous multi-agent middle layer (HMML) serving as a backbone module. HMML facilitates feature interaction across different levels, allowing for the exploration of richer features and enhancing the system’s detection performance. To mitigate the impact of spatiotemporal asynchrony on the results, we introduce the spatiotemporal asynchronous fusion (SAF) method. This approach involves learning complex nonlinear mapping relationships between input sequences and corresponding object sequences, enabling spatiotemporal alignment. Experimental validations conducted by V2ICooper on real-world DAIR-V2X-C dataset demonstrate superior accuracy and robustness in object detection. Additionally, the successful implementation of the proposed system in real scenarios substantiates its effectiveness, as evidenced by experimental results.