A distributed cooperative decision-making approach for multi-priority CAVs at unsignalized T-intersections
摘要
Current research on multi-agent reinforcement learning (MARL) for optimizing the passage of connected and autonomous vehicles (CAVs) predominantly focuses on regular four-way intersections, often neglecting the heterogeneity of CAVs in terms of access rights and scheduling priorities. To address this gap, this paper targets unsignalized T-intersections and systematically investigates the optimization of multi-priority CAVs passage strategies. A distributed MARL framework is proposed, integrating priority awareness and a multi-dimensional reward mechanism. Priority information is embedded into the observation space via one-hot encoding, enabling agents to recognize social roles and learn differentiated strategies. A hierarchical reward function—incorporating basic driving, safe avoidance, neighbor cooperation, and priority courtesy—is designed to improve coordination efficiency and rule adaptability. A PPO-based distributed training structure is constructed to enhance convergence stability. Computational and communication complexity and scalability are analyzed, including distributed training throughput, latency, and parallel efficiency. Extensive experiments on the MetaDrive platform, compared with SAC, DDPG, and FIFO baselines, demonstrate that the proposed method significantly outperforms existing approaches in key metrics such as passage efficiency, safety, and priority responsiveness. These results validate the effectiveness of the “multi-priority mechanism + multi-dimensional reward structure” in guiding the evolution of CAV passage strategies and provide a feasible approach for intelligent decision-making in complex traffic environments.