Currently, the relationship between human pose and energy expenditure in sports scenarios remains insufficiently explored, and traditional energy expenditure estimation methods face challenges in balancing accuracy and usability. To address this challenge, we propose a method for estimating energy expenditure based on human pose, effectively bridging this gap. Specifically, we propose a novel Spatio-temporal Pose Embedding-Enhanced Transformer (STPEFormer), designed to overcome the limitations of existing skeleton-based energy expenditure estimation methods in capturing the intricate patterns of human motion. The framework incorporates pose graph embeddings to enhance the features extracted by attention modules, introducing the Spatial Pose Embedding-Enhanced Attention (S-PEA) and Temporal Pose Embedding-Enhanced Attention (T-PEA) mechanisms. These modules exploit the skeletal topology and motion trajectories to inject localized spatiotemporal dependencies into the feature representations, which are further refined through a multi-head self-attention mechanism to capture high-relevance pose-to-metabolism patterns. By integrating these graph-based enhancements, STPEFormer effectively models spatial and temporal pose features in a unified and efficient manner. Extensive experiments on two benchmark datasets demonstrate the superior performance of STPEFormer in energy expenditure estimation. Specifically, the model outperforms all baselines on the E3V-K5 dataset and achieves performance comparable to video-based methods on the Vid2Burn-ADL dataset. Furthermore, the proposed graph-enhanced attention modules are fully compatible with existing Transformer-based architectures, highlighting their versatility and effectiveness.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

STPEFormer: Spatio-Temporal Pose Embedding-Enhanced Transformer for Energy Expenditure Estimation

  • Zhongteng Zhang,
  • Qing Peng,
  • Liu Zhang,
  • Zihao Zhang,
  • Jing Chong,
  • Weihong Huang

摘要

Currently, the relationship between human pose and energy expenditure in sports scenarios remains insufficiently explored, and traditional energy expenditure estimation methods face challenges in balancing accuracy and usability. To address this challenge, we propose a method for estimating energy expenditure based on human pose, effectively bridging this gap. Specifically, we propose a novel Spatio-temporal Pose Embedding-Enhanced Transformer (STPEFormer), designed to overcome the limitations of existing skeleton-based energy expenditure estimation methods in capturing the intricate patterns of human motion. The framework incorporates pose graph embeddings to enhance the features extracted by attention modules, introducing the Spatial Pose Embedding-Enhanced Attention (S-PEA) and Temporal Pose Embedding-Enhanced Attention (T-PEA) mechanisms. These modules exploit the skeletal topology and motion trajectories to inject localized spatiotemporal dependencies into the feature representations, which are further refined through a multi-head self-attention mechanism to capture high-relevance pose-to-metabolism patterns. By integrating these graph-based enhancements, STPEFormer effectively models spatial and temporal pose features in a unified and efficient manner. Extensive experiments on two benchmark datasets demonstrate the superior performance of STPEFormer in energy expenditure estimation. Specifically, the model outperforms all baselines on the E3V-K5 dataset and achieves performance comparable to video-based methods on the Vid2Burn-ADL dataset. Furthermore, the proposed graph-enhanced attention modules are fully compatible with existing Transformer-based architectures, highlighting their versatility and effectiveness.