Depth Decoupling for Bottom-Up Multi-Person 3D Pose Estimation
摘要
Recovering multi-person 3D poses from a single image is challenging due to depth ambiguity. Existing bottom-up methods show potential in resolving this by leveraging global contextual cues but suffer from unrelated information at two levels. At the task level, despite the heterogeneity between root and root-relative depth, these methods share information across different depth regressions, leading to negative transfer. At the information level, they treat the entire image equally during depth regression, introducing noise from irrelevant regions and hindering the acquisition of depth-related information. Both types of irrelevant information exacerbate the challenge of mitigating depth ambiguity. We propose a novel bottom-up approach, Depth Decoupling Network (DDNet), integrating Depth-related Decoupling Module (DRDM) and Depth-related Attention Module (DRAM). DRDM decouples feature subspaces of different depths and learns exclusive features for each depth by adaptively recalibrating responses and regulating latent variables distribution. DRAM introduces specific attention mechanisms for each depth regression to focus on task-related regions and mitigate noise from different perspectives, highlighting salient information spatially and channel-wise. Experiments on MuPoTS-3D and CMU Panoptic benchmarks show our method outperforms state-of-the-art bottom-up methods on \(\textrm{PCK}_{\text {rel}}\) , \(\textrm{PCK}_{\text {abs}}\) and MPJPE by at least 1.0%, 3.7% and 0.9mm.