<p>The application of multi-camera systems has enabled the capture of basketball players’ dynamic poses, providing valuable data support for coaches and analysts. However, numerous challenges still exist in practical applications. Firstly, in complex basketball game scenarios, particularly when players are moving quickly and experiencing occlusion, accurately extracting player poses remains highly challenging. Secondly, current methods lack effective spatio-temporal feature fusion when processing multi-camera data, making it difficult to fully capture the dynamic relationships between different viewpoints. This deficiency results in insufficient accuracy and stability in pose estimation. To address these issues, this paper proposes a multi-task learning model that combines local feature enhancement and multi-camera spatio-temporal feature fusion (LFEMSFF). First, the Local Feature Enhancement module utilizes a graph convolutional network (GCN) to extract detailed local player movements, such as limb bending and joint angles. This improves the model’s ability to capture complex pose variations and enhances its adaptability to occlusions and rapid movements. Next, the multi-camera spatio-temporal feature fusion module integrates data from different viewpoints using a spatio-temporal transformer network. This module considers not only the spatial information from each viewpoint but also leverages temporal sequence relationships to enhance the spatio-temporal continuity and resolution of the data, thereby capturing the dynamic changes in player movements more effectively. Finally, the multi-task learning module integrates the outputs from the previous two modules, optimizing both pose classification and keypoint localization tasks to ensure the accuracy and robustness of pose prediction. Experimental results show that the proposed model significantly improves pose detection accuracy and stability compared to existing methods on the NBA2K, Human3.6M and Sports-1 datasets.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhanced basketball pose estimation with spatio-temporal fusion and local feature learning

  • Wenyue Liu,
  • Zhihao Zhang,
  • Jianguo Qiu

摘要

The application of multi-camera systems has enabled the capture of basketball players’ dynamic poses, providing valuable data support for coaches and analysts. However, numerous challenges still exist in practical applications. Firstly, in complex basketball game scenarios, particularly when players are moving quickly and experiencing occlusion, accurately extracting player poses remains highly challenging. Secondly, current methods lack effective spatio-temporal feature fusion when processing multi-camera data, making it difficult to fully capture the dynamic relationships between different viewpoints. This deficiency results in insufficient accuracy and stability in pose estimation. To address these issues, this paper proposes a multi-task learning model that combines local feature enhancement and multi-camera spatio-temporal feature fusion (LFEMSFF). First, the Local Feature Enhancement module utilizes a graph convolutional network (GCN) to extract detailed local player movements, such as limb bending and joint angles. This improves the model’s ability to capture complex pose variations and enhances its adaptability to occlusions and rapid movements. Next, the multi-camera spatio-temporal feature fusion module integrates data from different viewpoints using a spatio-temporal transformer network. This module considers not only the spatial information from each viewpoint but also leverages temporal sequence relationships to enhance the spatio-temporal continuity and resolution of the data, thereby capturing the dynamic changes in player movements more effectively. Finally, the multi-task learning module integrates the outputs from the previous two modules, optimizing both pose classification and keypoint localization tasks to ensure the accuracy and robustness of pose prediction. Experimental results show that the proposed model significantly improves pose detection accuracy and stability compared to existing methods on the NBA2K, Human3.6M and Sports-1 datasets.