A Deep Reinforcement Learning Approach for Bipedal Robots Based on a Guiding Channel
摘要
To address the challenge of adapting to environmental nonlinearities, deep reinforcement learning (DRL) based on the Actor-Critic (AC) network architecture has become a widely adopted strategy in the simulation and practice of bipedal robot locomotion. Existing DRL methods based on the AC network architecture often exhibit a “Mount Everest phenomenon,” where the learning performance initially increases gradually but then suddenly drops at some stages. To tackle this issue, this paper primarily explores the stability issues in the learning process of deep reinforcement learning based on the introduction of a guiding channel within the AC network architecture. This is mainly dependent on the continuous gait library generated after traditional kinematic modeling and the application of a decaying guiding channel. Additionally, the current mainstream bipedal robot control algorithms based on the AC network architecture, such as the PPO and DDPG algorithms, are compared with our designed algorithm to mitigate the “Mount Everest effect”. The paper also contrasts different decay cutoff rounds and decay functions, and through an analysis of the loss function, it delves deeper into the role of the guiding channel.