Bootstrap Wi-Fi’s Own Latent: Towards Cross-Environment 3D Human Pose Estimation Using Wi-Fi
摘要
With advances in wireless technology and the popularity of wireless local area networks, wireless sensing-based pose estimation technology overcomes the privacy issues of traditional vision-based human perception solutions, demonstrating the potential in human-computer interaction, smart homes, healthcare, etc. However, achieving robust pose estimation using wireless sensing is challenging due to the diverse and dynamic usage environment of these applications. Furthermore, the low-resolution data collected by commercial WiFi devices is insufficient for the complex regression task of pose estimation, particularly for directly estimating joint position coordinates, leading to performance degradation. In this work, we introduce a cross-environment human pose estimation method using Wi-Fi’s Channel State Information (CSI) called BWOL. First, It employs a Divide-and-Conquer strategy, decoupling the pose estimation task into two simpler sub-tasks: Bone Length Estimation and Joint Angle Estimation. Thus, this strategy reduces the task complexity, ensuring high-quality pose estimation even with low-resolution data. Second, to achieve cross-environment generalization, BWOL uses a pre-trained environmental feature encoder to capture specific environmental features, assisting BWOL in completing the two subtasks in different environments. Finally, we combine the two subtasks using forward kinematics to achieve cross-environment posture estimation. We conducted extensive experiments on the MM-Fi dataset across four environments to demonstrate that BWOL significantly improves pose estimation accuracy and robustness, achieving an Mean Per Joint Position Error (MPJPE) of 130.43 mm in the basic scenario (32.90% improvement over the baseline) and 138.74 mm in the cross-environment (60.84% improvement over the baseline).