AGV Path Planning for Transmission Assembly Line Based on IRPA-PDERL Algorithm
摘要
Aiming at the issue of low path planning efficiency for Automated Guided Vehicle (AGV) facing mixed U-shaped and dynamic obstacles, an adaptive proximal distillation evolutionary reinforcement learning algorithm guided by intrinsic reward policy (IRPA-PDERL). Initially, Random Network Distillation is introduced as intrinsic rewards in the fitness function, which is enhanced the policy of diversity of elite policy evaluation. Secondly, an adaptive intrinsic reward weight factor \(\alpha\) is designed to balance the algorithm on the capacity to exploration and exploitation, aiding AGV in selecting the optimal strategy in highly uncertain dynamic obstacle environments. The decreased parameter sensitivity results from optimizing the covariance parameters of the mutation operator by covariance matrix adaption evolution strategy. Comparing with multiple algorithms on three different environments, the experimental results show that the improved algorithm reduces paths by 12.65%, 13.44%, and 12.87% in U-shaped sub-assembly line, dynamic obstacle, and transmission assembly line environments, respectively, with faster convergence speeds than the original algorithms, and has strong robustness in complex environments.