When discussing the 6D pose estimation problem of a single object, relying solely on feature points extracted from RGB data limits the perception capabilities of the shape and position of objects in three-dimensional space, leading to inferior accuracy during the registration process. Therefore, it is important to handle both the 2D image and the 3D geometric information in RGB-D data for 6D pose estimation. Existing RGB-D data fusion methods just stitch the data spatially without taking into account deep semantic integration of disparate modalities, failing to extract complementary informations. Because of this, the performance drops in scenarios where there are occlusions and changes in illumination. To address these issues, we present in this research a novel RGB-D processing fusion network for the 6D pose estimation. For the feature fusion module, we combine Hierarchical Dual Cross-modality Prompts (HDCP) with the original modality features, which are input into the subsequent processing network to achieve cross-spatial complementarity of different modalities. In addition, we use the modality similarity loss to measure their similarities in the feature space so that data from various modalities can be combined harmoniously. Extensive experiments demonstrate the effectiveness of the proposed modules.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

HDCP: Hierarchical Dual Cross-Modality Prompts Guided RGB-D Fusion for 6D Object Pose Estimation

  • Hanxue Fu,
  • Qiangchang Wang

摘要

When discussing the 6D pose estimation problem of a single object, relying solely on feature points extracted from RGB data limits the perception capabilities of the shape and position of objects in three-dimensional space, leading to inferior accuracy during the registration process. Therefore, it is important to handle both the 2D image and the 3D geometric information in RGB-D data for 6D pose estimation. Existing RGB-D data fusion methods just stitch the data spatially without taking into account deep semantic integration of disparate modalities, failing to extract complementary informations. Because of this, the performance drops in scenarios where there are occlusions and changes in illumination. To address these issues, we present in this research a novel RGB-D processing fusion network for the 6D pose estimation. For the feature fusion module, we combine Hierarchical Dual Cross-modality Prompts (HDCP) with the original modality features, which are input into the subsequent processing network to achieve cross-spatial complementarity of different modalities. In addition, we use the modality similarity loss to measure their similarities in the feature space so that data from various modalities can be combined harmoniously. Extensive experiments demonstrate the effectiveness of the proposed modules.