This paper presents a novel method for learning large scale vehicle navigation preference. Despite the recent increasing attention to preference learning with preference alignment methods such as Reinforcement Learning from Human Feedback (RLHF) or Inverse Reinforcement Learning (IRL), it is still challenging due to the lack of trajectory data. Because trajectory data representing agent behavior is essential for these methods. In this paper, we significantly reduces the dependence on trajectory data. Our method uses a Dual Deep Reinforcement Learning (D2RL) framework to simultaneously generate navigation preference data and learn the navigation policy based on the generated preference data for large scale vehicles. By designing detailed metrics, it is verified that the simulation results of large scale vehicle navigation process based on generated preferences match well with those guided by ground truth navigation preferences.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Large-Scale Vehicle Navigation Preference Generation Based on Dual Deep Reinforcement Learning

  • Yanshu Shuai,
  • Jiaoling Zheng,
  • Xingyu Fan,
  • Haoquan Wang,
  • Yu Zeng,
  • Peng Guo

摘要

This paper presents a novel method for learning large scale vehicle navigation preference. Despite the recent increasing attention to preference learning with preference alignment methods such as Reinforcement Learning from Human Feedback (RLHF) or Inverse Reinforcement Learning (IRL), it is still challenging due to the lack of trajectory data. Because trajectory data representing agent behavior is essential for these methods. In this paper, we significantly reduces the dependence on trajectory data. Our method uses a Dual Deep Reinforcement Learning (D2RL) framework to simultaneously generate navigation preference data and learn the navigation policy based on the generated preference data for large scale vehicles. By designing detailed metrics, it is verified that the simulation results of large scale vehicle navigation process based on generated preferences match well with those guided by ground truth navigation preferences.