Safe reinforcement learning for vision-based robotic manipulation in human-centered environments
摘要
Autonomous systems performing object manipulation in human-robot collaboration scenarios face fundamental challenges in balancing adaptability and safety. We present a reinforcement learning framework that addresses these challenges through safety-aware policy learning. Building on OpenAI’s Safety Gym, we extend its capabilities with a robotic arm model for practical manipulation tasks. Our approach employs end-to-end policy learning, comparing a constrained Lagrangian variant of Proximal Policy Optimization (cPPO) against standard PPO and Soft Actor-Critic (SAC) baselines. To handle high-dimensional visual inputs, we introduce a structured object-centric representation that captures multiple skills, objects, and their interactions, enabling goal-conditioned manipulation across diverse configurations. The agent trained on simple two-cube tasks generalizes to scenarios with six distinct objects in cluttered environments, demonstrating strong compositional generalization. Experimental results show that cPPO achieves superior safety performance, with an average episode cost of 15.26 compared to 18.03 for PPO and 19.48 for SAC. While cPPO’s task performance (average reward