MTFT: Multimodal Temporal Fusion Transformer for 3D Human Pose and Shape Estimation
摘要
Accurately and smoothly recovering 3D human motion from a monocular video is very challenging since minor parametric deviation can lead to a noticeable misalignment and contiguity between the estimated mesh and the input video. Existing methods of 3D human Pose and Shape Estimation (PSE) based on video typically adopt features extracted from multi-frame images to recover human mesh. The high complexity and low representation ability of these method often result in inconsistent pose motion and incorrect joint position alignment. In this paper, we aim to find a solution to align image features with skeleton features and to enhance the contiguity of inter-frame estimation. To this end, we propose a Multimodal Temporal Fusion Transformer (MTFT) for 3D human pose and mesh estimation. In MTFT, the Bidirectional Multimodal Fusion Module (BiMMFM) fuse the features extracted from images and skeleton features extracted from pose to align the estimated shape especially in the joints. The Multimodal Temporal Fusion Module (MMTFM) construct the temporal relationships between frames to enhance the contiguity of human pose and shape. In addition, we design a new fusion strategy Modal Softmax (MS) to fuse multimodal features and multi-frames features. The proposed MTFT is evaluated on three challenging benchmark datasets 3DPW, Human3.6M, and MPI-INF-3DHP, and achieves state-of-the-art results.