<p>Existing monocular video-based 3D pose estimation techniques often fail to fully exploit dynamic information across consecutive frames, resulting in incoherent pose estimations and significant degradation in joint localization accuracy under rapid motions or occlusions. To address these challenges, this paper proposes a novel Attention-guided Adjacent Frame-aware Network (AAFN) for 3D human pose estimation. The AAFN incorporates a Guided Temporal Attention Module (GTAM) to capture temporal dependencies and an External Guided Temporal Focus Module (EGTF) to enhance keyframe features. Based on the GTAM framework, we design a Neighborhood-Aware Temporal Encoder (NATE) to fuse local details from adjacent frames, coupled with bidirectional GRUs for feature smoothing, thereby constructing refined spatiotemporal representations. Experimental results demonstrate that the AAFN achieves MPJPE and PA-MPJPE scores of 63.8 mm and 42.2 mm, respectively, on the Human3.6M dataset, outperforming the baseline model TCMR by 13.3% and 18.8%.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Research on 3D human pose estimation via attention-guided adjacent frame-aware network

  • Jianwei Li,
  • Xuanyao Song,
  • Weitian Wang,
  • Chengxi Jiang,
  • Ziyi Zhuang

摘要

Existing monocular video-based 3D pose estimation techniques often fail to fully exploit dynamic information across consecutive frames, resulting in incoherent pose estimations and significant degradation in joint localization accuracy under rapid motions or occlusions. To address these challenges, this paper proposes a novel Attention-guided Adjacent Frame-aware Network (AAFN) for 3D human pose estimation. The AAFN incorporates a Guided Temporal Attention Module (GTAM) to capture temporal dependencies and an External Guided Temporal Focus Module (EGTF) to enhance keyframe features. Based on the GTAM framework, we design a Neighborhood-Aware Temporal Encoder (NATE) to fuse local details from adjacent frames, coupled with bidirectional GRUs for feature smoothing, thereby constructing refined spatiotemporal representations. Experimental results demonstrate that the AAFN achieves MPJPE and PA-MPJPE scores of 63.8 mm and 42.2 mm, respectively, on the Human3.6M dataset, outperforming the baseline model TCMR by 13.3% and 18.8%.