<p>With the shift towards online learning and autonomous training in physical education, Action Quality Assessment (AQA) has emerged as a crucial component in creating a closed-loop learning system. Existing methods, while improving accuracy and reliability, often lack interpretability. This paper proposes an interpretable and reliable AQA framework comprising two stages. First, a Motion information Enhanced Transformer (MiE Transformer) is introduced for 3D human pose estimation. By integrating Graph Convolutional Networks (GCNs) with Transformers, the MiE Transformer enhances action detail representation and motion dynamics, reducing joint motion errors through motion constraints in the loss function. We evaluate our model on popular benchmark dataset Human3.6M, the acceleration error was reduced to 1.0&#xa0;mm, and the joint position reconstruction also outperformed those of previous methods, with P1 error of 43.7&#xa0;mm. Second, the 3D pose data are decomposed into static and dynamic features, such as velocity, center of gravity changes, and normalized bone vectors. An improved Dynamic Time Warping (DTW) algorithm is then applied to quantify multidimensional differences between practice and standard actions. The final action quality score integrates multiple feature scores, demonstrating strong generalization and reliability in evaluating martial arts like TaiChi across varying body sizes and skill levels. This work not only advances the field of AQA but also highlights its potential for broad applications in online physical education.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Interpretable two-stage action quality assessment via 3D human pose estimation and dynamic feature alignment

  • Shouming Hou,
  • Aoyu Xia,
  • Zixuan Lu,
  • Weibo Yang

摘要

With the shift towards online learning and autonomous training in physical education, Action Quality Assessment (AQA) has emerged as a crucial component in creating a closed-loop learning system. Existing methods, while improving accuracy and reliability, often lack interpretability. This paper proposes an interpretable and reliable AQA framework comprising two stages. First, a Motion information Enhanced Transformer (MiE Transformer) is introduced for 3D human pose estimation. By integrating Graph Convolutional Networks (GCNs) with Transformers, the MiE Transformer enhances action detail representation and motion dynamics, reducing joint motion errors through motion constraints in the loss function. We evaluate our model on popular benchmark dataset Human3.6M, the acceleration error was reduced to 1.0 mm, and the joint position reconstruction also outperformed those of previous methods, with P1 error of 43.7 mm. Second, the 3D pose data are decomposed into static and dynamic features, such as velocity, center of gravity changes, and normalized bone vectors. An improved Dynamic Time Warping (DTW) algorithm is then applied to quantify multidimensional differences between practice and standard actions. The final action quality score integrates multiple feature scores, demonstrating strong generalization and reliability in evaluating martial arts like TaiChi across varying body sizes and skill levels. This work not only advances the field of AQA but also highlights its potential for broad applications in online physical education.