Enhancing Zero-Shot Skeleton-Based Action Recognition with Multi-semantic Action Descriptions
摘要
Zero-shot skeleton-based action recognition (ZSSAR) involves learning cross-modal alignment between visual and textual modalities, addressing the challenge of bridging the semantic gap between two modalities. This paper proposes a novel alignment method leveraging multi-semantic action descriptions (MSD) to enhance zero-shot recognition capabilities. Firstly, the MSD with rich semantic information, including action labels, the definition of action labels, and the execution processes of motion and human-object interactions, is constructed using GPT-4 to bridge the gap between textual and visual modalities. Secondly, an adaptive visual layer normalization module (ALN) processes visual features to minimize feature distribution disparities across categories and enhance model robustness and generalization. Finally, a contrastive learning method is used to achieve cross-modal alignment between two modalities using the MSD as an anchor point, effectively learning the distribution of visual features across different categories. Extensive experimental evaluations on the NTU RGB+D and NTU RGB+D 120 datasets demonstrate significant performance improvements, confirming the efficacy of our method for ZSSAR.