Virtual Command Allocation: Enhancing Hexapod Robot Locomotion Through Goal-Conditioned Reinforcement Learning
摘要
This study explores control methods for hexapod robots using reinforcement learning techniques. Goal-conditioned reinforcement learning (GCRL) is employed to develop a walking controller that allows smooth movement in any direction based on given commands. The high-dimensional action space of the hexapod robot often results in suboptimal policies. To address this issue, we propose the Virtual Command Allocation (VCA) algorithm, which enhances learning efficiency by introducing diverse training signals and avoiding common pitfalls in GCRL. Although GCRL problems can be addressed using standard reinforcement learning algorithms, these approaches often get trapped in local optima, which limits their effectiveness. To overcome this issue, we propose a new algorithm called Virtual Command Allocation (VCA). VCA is a straightforward method that involves modifying components of baseline reinforcement learning algorithms. Despite its simplicity, our validation experiments demonstrate that VCA can learn effectively without getting trapped in local optima, outperforming the baseline. Additionally, the proposed method has been shown to function well in parallelized environments, highlighting its scalability and practicality.