In the context of the rapid evolution of autonomous driving technology, the deployment of autonomous vehicle systems has witnessed significant expansion, underscoring the paramount importance of their path planning capabilities. This study proposes an improved Q-learning, specifically designed for the path planning of autonomous vehicles. Initially, the algorithm is enhanced with an advanced artificial potential field method for initializing the Q-table, thereby integrating prior knowledge and increasing the efficiency of the initial iterations. Further, the integration of eligibility traces as a memory mechanism within the algorithm ensures the prolonged influence of state-action pairs throughout the iterative process, thus boosting learning efficiency. Additionally, this paper introduces a dynamic λ adjustment strategy, tailored to meet the exploratory demands of autonomous vehicles and based on the simulated annealing algorithm. This strategy dynamically modulates the decay factor λ in accordance with the frequency of successful arrivals at the destinations, promoting more effective adaptation to varying environmental conditions. The empirical findings confirm the efficacy of the proposed Q-learning, particularly highlighting its utility in the path planning for autonomous vehicles.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Autonomous Vehicle Path Planning Algorithm Based on Improved Q-learning

  • Yuelong Wang,
  • Songyan Wang,
  • Tao Chao

摘要

In the context of the rapid evolution of autonomous driving technology, the deployment of autonomous vehicle systems has witnessed significant expansion, underscoring the paramount importance of their path planning capabilities. This study proposes an improved Q-learning, specifically designed for the path planning of autonomous vehicles. Initially, the algorithm is enhanced with an advanced artificial potential field method for initializing the Q-table, thereby integrating prior knowledge and increasing the efficiency of the initial iterations. Further, the integration of eligibility traces as a memory mechanism within the algorithm ensures the prolonged influence of state-action pairs throughout the iterative process, thus boosting learning efficiency. Additionally, this paper introduces a dynamic λ adjustment strategy, tailored to meet the exploratory demands of autonomous vehicles and based on the simulated annealing algorithm. This strategy dynamically modulates the decay factor λ in accordance with the frequency of successful arrivals at the destinations, promoting more effective adaptation to varying environmental conditions. The empirical findings confirm the efficacy of the proposed Q-learning, particularly highlighting its utility in the path planning for autonomous vehicles.