Joint Multi-person Body Detection and Orientation Estimation Via One Unified Embedding
摘要
Human body orientation estimation (HBOE) has been widely applied in various domains, including robotics, surveillance, and autonomous driving. Traditional approaches to HBOE typically assume that human instances are already identified and utilize well-cropped sub-images as input. However, such assumptions often prove to be invalid and inefficient in real-world scenarios involving crowds of people. To solve this problem, we propose a single-stage end-to-end trainable framework that estimates the locations and orientations of all bodies simultaneously by integrating bounding box prediction and direction angles into one unified embedding. Moreover, joint learning of HBOE and body detection integrates human features into body orientation estimation, improving the performance of HBOE, particularly in crowded and occluded scenarios. Extensive experiments on the reconstructed MEBOW dataset with more real scenarios demonstrate the effectiveness and efficiency of our approach. The code will be released.