Human Pose Estimation Based on Multiscale and Multilevel Feature Fusion Method
摘要
Due to changes in scale during feature extraction, human pose estimation networks often experience a loss of positional data, which compromises their precision. In order to improve human pose estimation, this study investigated the use of multi-tier feature extraction with multi-scale feature fusion in both spatial and channel dimensions. The hourglass network was used as the backbone, with multi-tiered features repeatedly employed from top-down to bottom-up to capture the spatial coordinates of each joint point. By integrating Atrous Spatial Pyramid Pooling (ASPP) and concurrent spatial and channel squeeze & excitation (scSE), a hybrid-atrous convolutional pooling pyramid network was constructed to facilitate multi-scale feature extraction and fusion, thereby minimizing the loss of feature information. An optimized pose estimation method based on feature fusion was proposed. The experimental results on the MPII dataset demonstrate that the proposed method achieves an accuracy of 92.9%, thus validating its ability to address scale variations in pose estimation and enhance the accuracy of human pose estimation.