ACDD: Multi-traffic Participant Interactive Motion Prediction with Agent -Centric Scene Modeling and Dual-Layer Decoding
摘要
Interactive trajectory prediction plays an indispensable role in autonomous driving systems, acting as a key link between perception and decision-making. In complex urban environments, the dynamic interactions among heterogeneous traffic participants and intricate road geometries present significant challenges for accurate trajectory forecasting. To address these challenges, we propose ACDD, a learning-based multi-traffic participant interactive trajectory prediction model equipped with a dual-layer decoder framework. First, we introduce a traffic participant-centered scene modeling method that designs a weighted cost function considering the distance, heading angle, and velocity between different traffic participants to extract the surrounding traffic participants of interest for each traffic participant at the current frame. Based on the extracted neighbors and comprehensive kinematic features, an interaction mask-based mask matrix is computed to facilitate efficient interaction modeling. Based on this mask, we design a spatio-temporal interaction encoder that selectively models interactions only with relevant neighboring traffic participants, significantly reducing computational complexity. Furthermore, we employ a dual-layer decoder architecture guided by intent query vectors to perform secondary optimization of predicted trajectories, effectively mitigating the common issue of mode collapse in traditional single-shot generation methods. Extensive experiments on the Argoverse 2 dataset show that our model achieves competitive performance. Additional ablation studies further validate the effectiveness and rationality of each component within the proposed framework.