Self-supervised motion forecasting with local information interaction in autonomous driving
摘要
Motion forecasting presents significant challenges critical for ensuring the safety of autonomous driving systems. The accuracy of these forecasts relies heavily on factors such as map topology and the behaviors of vehicles and pedestrians. However, within vast datasets, certain features with unique properties, capable of enhancing representation generalization often remain hidden and overlooked. While self-supervised learning (SSL) has shown promise in uncovering such hidden features through pretext tasks, its application to motion forecasting remains underexplored. In this paper, we propose a novel self-supervised motion forecasting method that exploits the interaction of map topology and actors’ maneuvers within localized focal points to generate more informative and generalizable representations for forecasting task. Since intersections, characterized by intricate structures and frequent motion state changes among actors, serve as pivotal locations where the topology of the intersection map profoundly influences actors’ intentions to change course, we leverage this interplay by calculating map structure-based actors’ attributes, and actors’ maneuver-based map attributes. These attributes yield significant advantages for motion forecasting tasks. Experimentally, our proposed method outperforms the baseline on both the challenging large-scale Argoverse benchmark (Chang et al. 2019) and local test, which demonstrates the effectiveness of the fusion of cross-domain information in a local neighborhood.
Graphical abstract