<p>In this work, we propose a novel distributional reinforcement learning (RL) approach, Kullback-Leibler Divergence regularized Distributional Actor-Critic (KDAC), to simultaneously address two fundamental challenges in off-policy RL: value function overestimation bias and accumulated approximation errors. Unlike existing multi-critic solutions, clipping the high-value quantiles in the distributional value function enables KDAC to suppress overestimation with a single critic network. Additionally, we introduce Kullback-Leibler divergence regularization between current and previous policies to mitigate accumulated approximation errors. The proposed lightweight KDAC efficiently reduces overestimated values while accelerating and stabilizing the learning process, ultimately enhancing sample efficiency. We evaluate on several benchmark tasks with different levels of complexity, where KDAC demonstrates significant competitive advantage in learning capability, overestimation reduction and sample efficiency compared with various traditional and advanced RL baselines, expanding the potential for more complicated control scenarios.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Reducing the value function over-estimation by Kullback-Leibler divergence regularized distributional actor-critic

  • Mingrong Gong,
  • Zhengkun Yi,
  • Yidong Chen,
  • Huiyun Li,
  • Yunduan Cui

摘要

In this work, we propose a novel distributional reinforcement learning (RL) approach, Kullback-Leibler Divergence regularized Distributional Actor-Critic (KDAC), to simultaneously address two fundamental challenges in off-policy RL: value function overestimation bias and accumulated approximation errors. Unlike existing multi-critic solutions, clipping the high-value quantiles in the distributional value function enables KDAC to suppress overestimation with a single critic network. Additionally, we introduce Kullback-Leibler divergence regularization between current and previous policies to mitigate accumulated approximation errors. The proposed lightweight KDAC efficiently reduces overestimated values while accelerating and stabilizing the learning process, ultimately enhancing sample efficiency. We evaluate on several benchmark tasks with different levels of complexity, where KDAC demonstrates significant competitive advantage in learning capability, overestimation reduction and sample efficiency compared with various traditional and advanced RL baselines, expanding the potential for more complicated control scenarios.