Decoupled Estimation of Human Pose and Shape for ICH Performance Video Based on L-C-HRNet
摘要
Aiming at the problem that the complex background and the presence of costume occlusion in the Intangible Cultural Heritage(ICH) performance video lead to large motion error and shape error in the human reconstruction results, we proposes a decoupled estimation method of the Human pose and shape for ICH performance video based on L-C-HRNet. First, a feature vector is extracted for each image frame using the feature extraction module proposed in this paper. Then the feature vector of each frame is inputted into the improved GRU layer to get the corresponding potential feature vector, and the regressor T with iterative feedback is used to get the SMPL parameters corresponding to each frame. In order to improve the motion accuracy of 3D human reconstruction, the \(\beta \) are optimized using the CLIP language-vision foundation model. Finally, a motion discriminator is introduced, which forms a Generative Adversarial Network(GAN) with the previous SMPL parameter generation part. Experiments show that our proposed method achieves the best performance on 3DPW dataset compared to MPS-NET and PMCE.