<p>Autonomous systems performing object manipulation in human-robot collaboration scenarios face fundamental challenges in balancing adaptability and safety. We present a reinforcement learning framework that addresses these challenges through safety-aware policy learning. Building on OpenAI’s Safety Gym, we extend its capabilities with a robotic arm model for practical manipulation tasks. Our approach employs end-to-end policy learning, comparing a constrained Lagrangian variant of Proximal Policy Optimization (cPPO) against standard PPO and Soft Actor-Critic (SAC) baselines. To handle high-dimensional visual inputs, we introduce a structured object-centric representation that captures multiple skills, objects, and their interactions, enabling goal-conditioned manipulation across diverse configurations. The agent trained on simple two-cube tasks generalizes to scenarios with six distinct objects in cluttered environments, demonstrating strong compositional generalization. Experimental results show that cPPO achieves superior safety performance, with an average episode cost of 15.26 compared to 18.03 for PPO and 19.48 for SAC. While cPPO’s task performance (average reward <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\sim\)</EquationSource> </InlineEquation>30) is slightly below PPO (<InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(\sim\)</EquationSource> </InlineEquation>35), it substantially outperforms SAC (<InlineEquation ID="IEq3"> <EquationSource Format="TEX">\(\sim\)</EquationSource> </InlineEquation>12). All algorithms converge by around 200,000 environment steps, with cPPO rapidly achieving safety compliance while sustaining learning progress. Although our evaluation is simulation-based, we have access to a 7-DOF xMatePro7 robotic arm for future hardware validation. Preliminary tests of low-level manipulation primitives show promise for sim-to-real transfer, and we plan to validate the complete framework on this platform in future work. These findings highlight the effectiveness of safe reinforcement learning for autonomous manipulation, advancing the practical deployment of collaborative robotic systems.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Safe reinforcement learning for vision-based robotic manipulation in human-centered environments

  • Fawad Khan,
  • Wei Feng,
  • Zhiyong Wang,
  • Tianlun Huang,
  • Xiao Liu,
  • Yunduan Cui,
  • Weijun Wang

摘要

Autonomous systems performing object manipulation in human-robot collaboration scenarios face fundamental challenges in balancing adaptability and safety. We present a reinforcement learning framework that addresses these challenges through safety-aware policy learning. Building on OpenAI’s Safety Gym, we extend its capabilities with a robotic arm model for practical manipulation tasks. Our approach employs end-to-end policy learning, comparing a constrained Lagrangian variant of Proximal Policy Optimization (cPPO) against standard PPO and Soft Actor-Critic (SAC) baselines. To handle high-dimensional visual inputs, we introduce a structured object-centric representation that captures multiple skills, objects, and their interactions, enabling goal-conditioned manipulation across diverse configurations. The agent trained on simple two-cube tasks generalizes to scenarios with six distinct objects in cluttered environments, demonstrating strong compositional generalization. Experimental results show that cPPO achieves superior safety performance, with an average episode cost of 15.26 compared to 18.03 for PPO and 19.48 for SAC. While cPPO’s task performance (average reward \(\sim\) 30) is slightly below PPO ( \(\sim\) 35), it substantially outperforms SAC ( \(\sim\) 12). All algorithms converge by around 200,000 environment steps, with cPPO rapidly achieving safety compliance while sustaining learning progress. Although our evaluation is simulation-based, we have access to a 7-DOF xMatePro7 robotic arm for future hardware validation. Preliminary tests of low-level manipulation primitives show promise for sim-to-real transfer, and we plan to validate the complete framework on this platform in future work. These findings highlight the effectiveness of safe reinforcement learning for autonomous manipulation, advancing the practical deployment of collaborative robotic systems.