Reinforcement learning (RL) has demonstrated significant potential in autonomous driving. However, its low sample efficiency and heavy reliance on environmental feedback pose challenges in handling highly dynamic and uncertain traffic scenarios, especially in pixel-based settings. In this paper, we propose a two-stage continual RL framework with implicit generative replay to address dynamic requirements in complex traffic situations. Based on the shared structures with parameterized skills and abstract transitions, we design a latent fine-tuning mechanism that facilitates steady policy improvement and avoids local optima. Additionally, we integrate diffusion-based implicit generative replay to enhance online experiences in latent space, promoting thorough exploration and exploitation. We also implement a dynamically decayed sampling ratio to effectively blend training data, mitigating distribution shifts’ impact and managing time costs. Extensive validation on complex and dynamic driving tasks shows that our approach significantly surpasses previous methods in both learning efficiency and generalization.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Continual Reinforcement Learning with Implicit Generative Replay for Autonomous Driving

  • Qi Deng,
  • Ruyang Li,
  • Qifu Hu,
  • Tengfei Zhang,
  • Heng Zhang

摘要

Reinforcement learning (RL) has demonstrated significant potential in autonomous driving. However, its low sample efficiency and heavy reliance on environmental feedback pose challenges in handling highly dynamic and uncertain traffic scenarios, especially in pixel-based settings. In this paper, we propose a two-stage continual RL framework with implicit generative replay to address dynamic requirements in complex traffic situations. Based on the shared structures with parameterized skills and abstract transitions, we design a latent fine-tuning mechanism that facilitates steady policy improvement and avoids local optima. Additionally, we integrate diffusion-based implicit generative replay to enhance online experiences in latent space, promoting thorough exploration and exploitation. We also implement a dynamically decayed sampling ratio to effectively blend training data, mitigating distribution shifts’ impact and managing time costs. Extensive validation on complex and dynamic driving tasks shows that our approach significantly surpasses previous methods in both learning efficiency and generalization.