Despite advances in continuous-time reinforcement learning, employing continuous-time actor-critic (CTAC) algorithm is not always the best choice due to the lack of value estimation method. This paper introduces a novel Continuous-time Double Actors and Regularized Critics method. We apply the underlying ideas behind the success of continuous control reinforcement learning to continuous-time tasks, aiming to achieve better value estimation and exploration compared to CTAC. Experimental results demonstrate that our method can reach the maximum reward in fewer training rounds, significantly outperforming CTAC.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Continuous-Time Double Actors and Regularized Critics in Reinforcement Learning

  • Shuhao Li,
  • Shengjie Zhang,
  • Han Zhang,
  • Rui Zhou

摘要

Despite advances in continuous-time reinforcement learning, employing continuous-time actor-critic (CTAC) algorithm is not always the best choice due to the lack of value estimation method. This paper introduces a novel Continuous-time Double Actors and Regularized Critics method. We apply the underlying ideas behind the success of continuous control reinforcement learning to continuous-time tasks, aiming to achieve better value estimation and exploration compared to CTAC. Experimental results demonstrate that our method can reach the maximum reward in fewer training rounds, significantly outperforming CTAC.