Neural Lyapunov-Guided Reinforcement Learning with LIPM Foothold Convergence for Stable Bipedal Locomotion
摘要
This paper proposes a Reinforcement Learning framework for bipedal robots that integrates Lyapunov stability principles with gait prediction based on the Linear Inverted Pendulum Model (LIPM). Unlike conventional approaches that rigidly enforce the policy to track model-planned footholds, the proposed method emphasizes the convergence trend whereby actual footholds gradually approach the desired ones. To achieve this, a neural Lyapunov critic is constructed to incorporate LIPM-predicted footholds and their errors as inputs, learning energy-decreasing patterns to softly constrain policy behavior. Furthermore, a Lyapunov-augmented advantage estimation mechanism is developed, enabling the policy to benefit from both task rewards and stability compensation, thereby unifying physical interpretability, stability guarantees, and the flexibility of Reinforcement Learning. This study conducts large-scale experiments in Isaac Gym with 4096 parallel environments and further validate the approach on the TRON1A bipedal robot. Results demonstrate that the proposed method achieves superior performance in terrain adaptability, velocity tracking capability, stair-climbing success rate, and highly dynamic locomotion tasks. Moreover, the approach exhibits strong sim-to-real transferability, confirming its robustness and generalization in diverse real-world environments.