Making up for the lack of generalization and environmental simulation of traditional algorithms, a motion planning method of space manipulator based on reinforcement learning is designed. First, the standard Denavit-Hartenberg(DH) model of space manipulator is given. Further, combined with the characteristics of the space mission, the state space, action space and reward functions are designed. The Proximal Policy Optimization(PPO) is used as the framework to realize the motion planning task of the space manipulator. ISAAC GYM is chosed as the simulation platform to improve the training speed and strategy generalization ability through the setting of multiagent training and environment randomization at the same time. The simulation results show that the proposed method can realize the task of grasping the object by avoiding obstacles in the case of space microgravity, and the method has strong practicability and effectiveness.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Research on Obstacle Avoidance Motion Planning of Space Manipulator Based on Reinforcement Learning

  • Zixuan Zhang,
  • Wei Dong,
  • Chunyan Wang,
  • Jing Sun

摘要

Making up for the lack of generalization and environmental simulation of traditional algorithms, a motion planning method of space manipulator based on reinforcement learning is designed. First, the standard Denavit-Hartenberg(DH) model of space manipulator is given. Further, combined with the characteristics of the space mission, the state space, action space and reward functions are designed. The Proximal Policy Optimization(PPO) is used as the framework to realize the motion planning task of the space manipulator. ISAAC GYM is chosed as the simulation platform to improve the training speed and strategy generalization ability through the setting of multiagent training and environment randomization at the same time. The simulation results show that the proposed method can realize the task of grasping the object by avoiding obstacles in the case of space microgravity, and the method has strong practicability and effectiveness.