<p>Existing task-oriented grasping approaches often rely on 2D pixel-wise affordance segmentation or predefined part annotations, limiting their applicability in unstructured 3D environments and constraining the grasp planning space. To overcome these limitations, we introduce a novel affordance-labeled grasp dataset constructed on simulation, capturing diverse functional interactions across object categories in a 6-DoF space. Building on this foundation, we propose a unified, language-guided grasping framework that takes partial point clouds and natural language instructions as input to generate semantically meaningful and geometrically feasible grasp poses. Specifically, a vision-language affordance grounding module produces dense 3D affordance maps aligned with task semantics, and a task-oriented grasp pipeline predicts coarse grasp candidates with implicit affordance cues. The coarse grasp proposals are subsequently refined based on visual affordance guidance, significantly enhancing both semantic alignment and grasp practicality. Extensive experiments in synthetic and real-world scenarios demonstrate that our method outperforms state-of-the-art approaches, effectively generalizing across diverse objects and tasks.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing task-oriented robotic grasping via 3D affordance grounding from vision-language models

  • Wenkai Chen,
  • Shang-Ching Liu,
  • Qingdu Li,
  • Yung-Hui Li,
  • Jianwei Zhang

摘要

Existing task-oriented grasping approaches often rely on 2D pixel-wise affordance segmentation or predefined part annotations, limiting their applicability in unstructured 3D environments and constraining the grasp planning space. To overcome these limitations, we introduce a novel affordance-labeled grasp dataset constructed on simulation, capturing diverse functional interactions across object categories in a 6-DoF space. Building on this foundation, we propose a unified, language-guided grasping framework that takes partial point clouds and natural language instructions as input to generate semantically meaningful and geometrically feasible grasp poses. Specifically, a vision-language affordance grounding module produces dense 3D affordance maps aligned with task semantics, and a task-oriented grasp pipeline predicts coarse grasp candidates with implicit affordance cues. The coarse grasp proposals are subsequently refined based on visual affordance guidance, significantly enhancing both semantic alignment and grasp practicality. Extensive experiments in synthetic and real-world scenarios demonstrate that our method outperforms state-of-the-art approaches, effectively generalizing across diverse objects and tasks.