Category-level object pose estimation plays a crucial role in a wide range of practical applications by accurately predicting the poses and sizes of unseen objects within a specific category. However, accurately estimating object poses remains a significant challenge due to substantial shape variations within the same category. To address this issue, this paper introduces a novel learning network for object pose estimation that is guided by a shape descriptor. By capturing the geometric information of an object’s shape, the shape descriptor provides valuable input for subsequent feature learning, effectively handling shape variations. Moreover, our framework incorporates a confidence-based pose estimator, which assigns confidence scores to each pose prediction. This integration allows for the acquisition of more accurate poses with higher confidence by penalizing poses with low confidence. Experimental results on the CAMERA25 and REAL275 datasets demonstrate the superiority of our approach over state-of-the-art methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Shape Descriptor Guided Learning for Category-Level Object Pose Estimation

  • Yun Liu,
  • Weiming Wang,
  • Fu Lee Wang,
  • Haoran Xie,
  • Honghua Chen,
  • Mingqiang Wei,
  • Jing Qin

摘要

Category-level object pose estimation plays a crucial role in a wide range of practical applications by accurately predicting the poses and sizes of unseen objects within a specific category. However, accurately estimating object poses remains a significant challenge due to substantial shape variations within the same category. To address this issue, this paper introduces a novel learning network for object pose estimation that is guided by a shape descriptor. By capturing the geometric information of an object’s shape, the shape descriptor provides valuable input for subsequent feature learning, effectively handling shape variations. Moreover, our framework incorporates a confidence-based pose estimator, which assigns confidence scores to each pose prediction. This integration allows for the acquisition of more accurate poses with higher confidence by penalizing poses with low confidence. Experimental results on the CAMERA25 and REAL275 datasets demonstrate the superiority of our approach over state-of-the-art methods.