3D hand pose estimation from monocular RGB images is an essential topic in computer vision and pattern recognition, which is widely used in various fields, especially for virtual reality, human-computer interaction, and gesture recognition, but often grapples with challenges such as pose complexity and self-occlusions. Most existing methods often fail to sufficiently capture the skeletal representation of the hand due to the complex topology relationships between hand joints. To mitigate these challenges, we introduce the Multilevel Topology Structure-aware Network (MTS-Net), a novel deep learning architecture designed for further exploring the hierarchy of hand topology to improve 3D hand pose estimation. Specifically, we propose a Multi-features Cross-Attention module (MCA) to enhance the interaction of multi-level information from different hierarchical topologies of joint, part, and hand. Finally, to validate the effectiveness of the proposed model, we conduct experimental verification on three commonly used public datasets: Rendered Hand Dataset (RHD), Stereo Hand Pose Benchmark (STB), and First-Person Hand Action Benchmark (FPHA). The experimental results surpassed those of the existing state-of-the-art models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multilevel Topology Structure-Aware Network for 3D Hand Pose Estimation

  • Yanjun Liu,
  • Wanshu Fan,
  • Xiaopeng Wei,
  • Dongsheng Zhou

摘要

3D hand pose estimation from monocular RGB images is an essential topic in computer vision and pattern recognition, which is widely used in various fields, especially for virtual reality, human-computer interaction, and gesture recognition, but often grapples with challenges such as pose complexity and self-occlusions. Most existing methods often fail to sufficiently capture the skeletal representation of the hand due to the complex topology relationships between hand joints. To mitigate these challenges, we introduce the Multilevel Topology Structure-aware Network (MTS-Net), a novel deep learning architecture designed for further exploring the hierarchy of hand topology to improve 3D hand pose estimation. Specifically, we propose a Multi-features Cross-Attention module (MCA) to enhance the interaction of multi-level information from different hierarchical topologies of joint, part, and hand. Finally, to validate the effectiveness of the proposed model, we conduct experimental verification on three commonly used public datasets: Rendered Hand Dataset (RHD), Stereo Hand Pose Benchmark (STB), and First-Person Hand Action Benchmark (FPHA). The experimental results surpassed those of the existing state-of-the-art models.