To address the challenge of adapting to environmental nonlinearities, deep reinforcement learning (DRL) based on the Actor-Critic (AC) network architecture has become a widely adopted strategy in the simulation and practice of bipedal robot locomotion. Existing DRL methods based on the AC network architecture often exhibit a “Mount Everest phenomenon,” where the learning performance initially increases gradually but then suddenly drops at some stages. To tackle this issue, this paper primarily explores the stability issues in the learning process of deep reinforcement learning based on the introduction of a guiding channel within the AC network architecture. This is mainly dependent on the continuous gait library generated after traditional kinematic modeling and the application of a decaying guiding channel. Additionally, the current mainstream bipedal robot control algorithms based on the AC network architecture, such as the PPO and DDPG algorithms, are compared with our designed algorithm to mitigate the “Mount Everest effect”. The paper also contrasts different decay cutoff rounds and decay functions, and through an analysis of the loss function, it delves deeper into the role of the guiding channel.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Deep Reinforcement Learning Approach for Bipedal Robots Based on a Guiding Channel

  • Jialei Wang,
  • Junxiao Li,
  • Shutong Zhang

摘要

To address the challenge of adapting to environmental nonlinearities, deep reinforcement learning (DRL) based on the Actor-Critic (AC) network architecture has become a widely adopted strategy in the simulation and practice of bipedal robot locomotion. Existing DRL methods based on the AC network architecture often exhibit a “Mount Everest phenomenon,” where the learning performance initially increases gradually but then suddenly drops at some stages. To tackle this issue, this paper primarily explores the stability issues in the learning process of deep reinforcement learning based on the introduction of a guiding channel within the AC network architecture. This is mainly dependent on the continuous gait library generated after traditional kinematic modeling and the application of a decaying guiding channel. Additionally, the current mainstream bipedal robot control algorithms based on the AC network architecture, such as the PPO and DDPG algorithms, are compared with our designed algorithm to mitigate the “Mount Everest effect”. The paper also contrasts different decay cutoff rounds and decay functions, and through an analysis of the loss function, it delves deeper into the role of the guiding channel.