<p>Diffusion Policy has emerged as a powerful approach for imitation learning in robotic manipulation. However, policies trained solely on RGB observations often fail to capture the spatial structure of a scene, which limits performance on tasks requiring precise distance-aware motions, such as approaching an object along the camera axis. A natural remedy is to provide depth as an additional input, but two practical challenges remain: designing a conditioning mechanism that encourages the policy to use depth effectively and obtaining reliable depth when hardware measurements are unavailable or noisy. This paper presents Depth-Augmented Diffusion Policy, a framework that addresses both issues. The method employs separate visual encoders for RGB images and depth maps and fuses their representations through stochastic feature-level addition with modality dropout, reducing modality dominance and promoting the use of geometric cues during action generation. In addition, the results demonstrate that pseudo-depth generated by the pre-trained Depth Anything V2 estimator can replace sensor depth, removing the dependence on an RGB-D camera. The proposed approach is evaluated on a real-world robotic manipulation task performed using the Somatic Jinx robot, in which the robot must close a toilet-seat lid. The proposed policy achieves a 95% success rate with sensor depth and 90% with pseudo-depth, compared with 50% for the RGB-only baseline and 75% for 3D Diffusion Policy. These results show that explicit depth conditioning substantially improves performance on this distance-sensitive manipulation task and that pseudo-depth can serve as an effective alternative to hardware depth in this setting. Additional evaluation is conducted on the LIBERO simulation benchmark. Pseudo-depth augmentation yields a modest but consistent improvement across three task suites. The results indicate that the benefit extends beyond the single real-world task but remains most pronounced in distance-sensitive manipulation. </p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Depth-augmented diffusion policy with pseudo-depth for robust robotic manipulation

  • Lesia Hentosh,
  • Yaroslav Hladun,
  • Ivan Tsmots,
  • Oleh Bisikalo

摘要

Diffusion Policy has emerged as a powerful approach for imitation learning in robotic manipulation. However, policies trained solely on RGB observations often fail to capture the spatial structure of a scene, which limits performance on tasks requiring precise distance-aware motions, such as approaching an object along the camera axis. A natural remedy is to provide depth as an additional input, but two practical challenges remain: designing a conditioning mechanism that encourages the policy to use depth effectively and obtaining reliable depth when hardware measurements are unavailable or noisy. This paper presents Depth-Augmented Diffusion Policy, a framework that addresses both issues. The method employs separate visual encoders for RGB images and depth maps and fuses their representations through stochastic feature-level addition with modality dropout, reducing modality dominance and promoting the use of geometric cues during action generation. In addition, the results demonstrate that pseudo-depth generated by the pre-trained Depth Anything V2 estimator can replace sensor depth, removing the dependence on an RGB-D camera. The proposed approach is evaluated on a real-world robotic manipulation task performed using the Somatic Jinx robot, in which the robot must close a toilet-seat lid. The proposed policy achieves a 95% success rate with sensor depth and 90% with pseudo-depth, compared with 50% for the RGB-only baseline and 75% for 3D Diffusion Policy. These results show that explicit depth conditioning substantially improves performance on this distance-sensitive manipulation task and that pseudo-depth can serve as an effective alternative to hardware depth in this setting. Additional evaluation is conducted on the LIBERO simulation benchmark. Pseudo-depth augmentation yields a modest but consistent improvement across three task suites. The results indicate that the benefit extends beyond the single real-world task but remains most pronounced in distance-sensitive manipulation.