The goal of 2D human pose estimation is to extract human key points from images or videos to obtain pose information. This technology has significant applications in areas such as action recognition, human-computer interaction, and sports training. With the continuous development of application fields, the requirements for high accuracy and real-time performance in 2D human pose estimation models have become increasingly stringent. To address these challenges, this paper proposes a lightweight 2D human pose estimation network based on multi-scale fusion and an attention mechanism. Considering that multi-level sampling and fusion during the up-sampling process often lead to insufficient interaction between mid-resolution features, and that low-resolution features suffer from loss of fine details, we introduce a channel-spatial cross-attention mechanism. By incorporating this mechanism into low-resolution feature channels, we effectively mitigate information loss during the up-sampling process. During the feature sampling stage, some resolution features need to undergo both up-sampling and down-sampling operations simultaneously to achieve effective cross-resolution feature fusion. To this end, we design a multi-resolution fusion module to supplement the missing details and contextual information in mid-resolution features. Experimental results demonstrate that on the COCO and MPII datasets, our method achieves significant improvements in key metrics such as AP (Average Precision) and AP50, while adding only a minimal increase in model complexity.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Lightweight 2D Human Pose Estimation Based on Multi-scale Fusion and Attention Mechanism

  • Boyu Qi,
  • Ling Wang,
  • Jing Dong,
  • Pengfei Yi,
  • Xiaoyong Fang,
  • Rui Liu

摘要

The goal of 2D human pose estimation is to extract human key points from images or videos to obtain pose information. This technology has significant applications in areas such as action recognition, human-computer interaction, and sports training. With the continuous development of application fields, the requirements for high accuracy and real-time performance in 2D human pose estimation models have become increasingly stringent. To address these challenges, this paper proposes a lightweight 2D human pose estimation network based on multi-scale fusion and an attention mechanism. Considering that multi-level sampling and fusion during the up-sampling process often lead to insufficient interaction between mid-resolution features, and that low-resolution features suffer from loss of fine details, we introduce a channel-spatial cross-attention mechanism. By incorporating this mechanism into low-resolution feature channels, we effectively mitigate information loss during the up-sampling process. During the feature sampling stage, some resolution features need to undergo both up-sampling and down-sampling operations simultaneously to achieve effective cross-resolution feature fusion. To this end, we design a multi-resolution fusion module to supplement the missing details and contextual information in mid-resolution features. Experimental results demonstrate that on the COCO and MPII datasets, our method achieves significant improvements in key metrics such as AP (Average Precision) and AP50, while adding only a minimal increase in model complexity.