<p>Human pose estimation (HPE) is a crucial research direction in computer vision, offering non-invasive advantages in sports motion analysis with potential for deployment in intelligent coaching scenarios. Current state-of-the-art HPE algorithms predominantly employ task-specific Transformer architectures tailored for pose estimation. However, these methods still suffer from insufficient keypoint localization accuracy when handling scenarios involving human body occlusion, which is particularly pronounced during rapid badminton movements such as swinging and smashing. To address this, we leverage large-scale visual pre-training by fine-tuning the DINOv2 foundation model as a feature extractor, introducing a structure-aware collaborative mechanism to handle upper limb cross-occlusion. Extensive experiments on COCO, CrowdPose, and a custom badminton pose dataset demonstrate our method’s effectiveness, with D<InlineEquation ID="IEq4"> <EquationSource Format="TEX">\(^{2}\)</EquationSource> </InlineEquation>Pose-L achieving AP scores of 79.6 on COCO and 77.3 on CrowdPose, and showing notable advantages in challenging occlusion scenarios, providing quantitative support for intelligent badminton coaching systems.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

D\(^{2}\)pose: decoupling localization and structure for robust human pose estimation via DINOv2

  • Ruihan Zheng,
  • Ruijian Fang,
  • Fan Zhang

摘要

Human pose estimation (HPE) is a crucial research direction in computer vision, offering non-invasive advantages in sports motion analysis with potential for deployment in intelligent coaching scenarios. Current state-of-the-art HPE algorithms predominantly employ task-specific Transformer architectures tailored for pose estimation. However, these methods still suffer from insufficient keypoint localization accuracy when handling scenarios involving human body occlusion, which is particularly pronounced during rapid badminton movements such as swinging and smashing. To address this, we leverage large-scale visual pre-training by fine-tuning the DINOv2 foundation model as a feature extractor, introducing a structure-aware collaborative mechanism to handle upper limb cross-occlusion. Extensive experiments on COCO, CrowdPose, and a custom badminton pose dataset demonstrate our method’s effectiveness, with D \(^{2}\) Pose-L achieving AP scores of 79.6 on COCO and 77.3 on CrowdPose, and showing notable advantages in challenging occlusion scenarios, providing quantitative support for intelligent badminton coaching systems.