Large-Scale Vehicle Navigation Preference Generation Based on Dual Deep Reinforcement Learning
摘要
This paper presents a novel method for learning large scale vehicle navigation preference. Despite the recent increasing attention to preference learning with preference alignment methods such as Reinforcement Learning from Human Feedback (RLHF) or Inverse Reinforcement Learning (IRL), it is still challenging due to the lack of trajectory data. Because trajectory data representing agent behavior is essential for these methods. In this paper, we significantly reduces the dependence on trajectory data. Our method uses a Dual Deep Reinforcement Learning (D2RL) framework to simultaneously generate navigation preference data and learn the navigation policy based on the generated preference data for large scale vehicles. By designing detailed metrics, it is verified that the simulation results of large scale vehicle navigation process based on generated preferences match well with those guided by ground truth navigation preferences.