PoseRAC: Enhancing Repetitive Action Counting with Salient Poses
摘要
Repetitive action counting aims to count the number of periodic actions in a video, offering significant application values for human activities. However, this task has not been extensively explored, and previous methods struggle to address the periodic representation. Through an analysis of the relationship between human poses and actions, we present a novel concept called Salient Pose, which effectively represents each action. By further linking these salient poses and repetitive actions, we introduce a new approach for repetitive action counting called PoseRAC, to model the relationship between salient pose and actions and complete the counting task based on salient poses. Leveraging the foundation generative models, our model can perform zero-shot predictions without using any training set. Furthermore, by incorporating an off-the-shelf text encoder, our model can count unseen actions in an open-set setting. Our approach achieves state-of-the-art performance on three mainstream benchmarks: RepCount, UCFRep, and Countix.