<p>Robotics is reshaping industries and redefining work paradigms. This paper focuses on vision recognition systems and their critical role in enabling dexterous grasping for collaborative robots. While traditional vision systems often struggle with depth perception and target identification, binocular vision has emerged as a promising alternative, offering improved depth understanding and 3D reconstruction for more accurate and agile grasping. However, integrating binocular sensors for real-time, high-precision object detection and grasp planning in complex environments remains challenging. To address this, we first tackle 6D pose estimation using binocular vision. Conventional methods typically depend on precise object modeling, which limits their applicability. We propose a model-agnostic pose estimation network based on self-attention, overcoming these constraints. Evaluations on the LINEMOD dataset demonstrate that our method outperforms the established benchmark Gen6D in both ADD(-S) and 2D projection metrics. In parallel pinch grasping, traditional pose estimation methods often fail in cluttered scenarios. To overcome this, we introduce a novel 7-DoF grasping pose generation framework that integrates sparse attention and null convolution. This design enhances the model’s ability to capture fine-grained features from point clouds, significantly improving the accuracy of parallel gripping pose estimation. Extensive experiments on the GraspNet-1Billion dataset validate the effectiveness of our approach, especially in unseen object scenarios, where it achieves an improvement of 6.12 AP, demonstrating strong generalization capability.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Model-agnostic pose estimation for enhanced collaborative robot grasping via binocular vision

  • Hui Zhang,
  • Yue Wang,
  • Kang An,
  • Yijie Wang,
  • Ruoying Shi

摘要

Robotics is reshaping industries and redefining work paradigms. This paper focuses on vision recognition systems and their critical role in enabling dexterous grasping for collaborative robots. While traditional vision systems often struggle with depth perception and target identification, binocular vision has emerged as a promising alternative, offering improved depth understanding and 3D reconstruction for more accurate and agile grasping. However, integrating binocular sensors for real-time, high-precision object detection and grasp planning in complex environments remains challenging. To address this, we first tackle 6D pose estimation using binocular vision. Conventional methods typically depend on precise object modeling, which limits their applicability. We propose a model-agnostic pose estimation network based on self-attention, overcoming these constraints. Evaluations on the LINEMOD dataset demonstrate that our method outperforms the established benchmark Gen6D in both ADD(-S) and 2D projection metrics. In parallel pinch grasping, traditional pose estimation methods often fail in cluttered scenarios. To overcome this, we introduce a novel 7-DoF grasping pose generation framework that integrates sparse attention and null convolution. This design enhances the model’s ability to capture fine-grained features from point clouds, significantly improving the accuracy of parallel gripping pose estimation. Extensive experiments on the GraspNet-1Billion dataset validate the effectiveness of our approach, especially in unseen object scenarios, where it achieves an improvement of 6.12 AP, demonstrating strong generalization capability.