Visibility-Guided GCN-Transformer: Enhancing 2D Pose Estimation Under Occlusion
摘要
Occlusion remains a core challenge in human pose estimation. Because visual evidence is missing, accurate localization of hidden keypoints must exploit both local anatomical cues and global pose consistency. Yet the ternary visibility labels that quantify occlusion severity in mainstream datasets are often ignored. Accordingly, we propose ViGTNet, a visibility-guided GCN–Transformer alternating architecture that explicitly incorporates these labels for efficient local–global collaboration. ViGTNet first predicts keypoint visibility, then strengthens feature propagation from fully visible to occluded keypoints along graph edges and suppresses occlusion noise within global self-attention, while additionally emphasizing partially occluded keypoints during training via a dynamic loss-reweighting scheme. ViGTNet improves overall AP by 1.0 on COCO and APHard by 0.8 on CrowdPose, demonstrating robust performance under occlusion.