Non-prehensile manipulation plays a crucial role in the field of robotics, especially when objects are irregular, cumbersome or heavy. Traditional approaches are usually implemented based on reinforcement learning or imitation learning. However, reinforcement learning methods suffer from the low sampling efficiency with the environment, while imitation learning based ones rely on the time-consuming demonstrations collection. To this end, this paper proposes a novel Causal Policy Learning scheme from Self-play (CaPLS) for pushing manipulation. On one hand, CaPLS provides a light-weight way that the causal prior is learnt from the random self-play data and encoded as a structural causal model (SCM) using a graph neural network (GNN). On the other hand, SCM is used to guide the policy learning combining imitation learning and reinforcement learning. The demonstrations for imitation learning are collected using SCM directly, resulting in initially optimized policy. The reinforcement learning is then implemented to refine the policy for a specific task. It greatly reduces the iteration of interactions in this manner. Experimental results demonstrate that the proposed method can outperform the SOTA methods, and the generalization capability is improved as well. Code will be made publicly available.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Causal Policy Learning from Self-play for Effective Non-prehensile Planar Manipulation

  • Zepeng Sun,
  • Yong Guan,
  • Jie Zhang,
  • Zhiping Shi,
  • Zhenzhou Shao

摘要

Non-prehensile manipulation plays a crucial role in the field of robotics, especially when objects are irregular, cumbersome or heavy. Traditional approaches are usually implemented based on reinforcement learning or imitation learning. However, reinforcement learning methods suffer from the low sampling efficiency with the environment, while imitation learning based ones rely on the time-consuming demonstrations collection. To this end, this paper proposes a novel Causal Policy Learning scheme from Self-play (CaPLS) for pushing manipulation. On one hand, CaPLS provides a light-weight way that the causal prior is learnt from the random self-play data and encoded as a structural causal model (SCM) using a graph neural network (GNN). On the other hand, SCM is used to guide the policy learning combining imitation learning and reinforcement learning. The demonstrations for imitation learning are collected using SCM directly, resulting in initially optimized policy. The reinforcement learning is then implemented to refine the policy for a specific task. It greatly reduces the iteration of interactions in this manner. Experimental results demonstrate that the proposed method can outperform the SOTA methods, and the generalization capability is improved as well. Code will be made publicly available.