Development of depth sensors and pose estimation algorithms has sparked significant interest and a range of applications for human skeleton action recognition utilizing graph convolutional networks. Advanced techniques, which adeptly harness the complexities of joint, bone, and motion features, enable a dynamic documentation of the nuances in human movement, achieved through an array of topological transformations. However, many models struggle to distinguish actions that share similar trajectories. In an effort to overcome these hurdles, our Multi-level feature-based semantic guided neural network (MFSGN) pioneers the integration of fourth-order angular features, not only strengthening the recognition of joint and body-part connections but also providing a nuanced lens through which actions with similar motion trajectories can be more accurately differentiated. We explicitly introduce high-level semantics of nodes in the network (node types and frame indexes) to improve the feature representation of nodes. To significantly improve the spatial correlation essential for accurate data interpretation, we present a sophisticated spatial Data-driven excitation (SDE) module. Moreover, integrating a streamlined Multi-scale Temporal Convolution (MS-TCN) network serves to significantly strengthen the representation of temporal features. Our method has been meticulously validated against the NTU-RGB D 60 and NTU-RGB D 120 datasets, and the findings confirm its superiority over existing state-of-the-art methods in terms of accuracy and robustness.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-level Feature-Based Semantic Guided Neural Network for Skeleton Action Recognition

  • Yongfeng Qi,
  • Anye Liang

摘要

Development of depth sensors and pose estimation algorithms has sparked significant interest and a range of applications for human skeleton action recognition utilizing graph convolutional networks. Advanced techniques, which adeptly harness the complexities of joint, bone, and motion features, enable a dynamic documentation of the nuances in human movement, achieved through an array of topological transformations. However, many models struggle to distinguish actions that share similar trajectories. In an effort to overcome these hurdles, our Multi-level feature-based semantic guided neural network (MFSGN) pioneers the integration of fourth-order angular features, not only strengthening the recognition of joint and body-part connections but also providing a nuanced lens through which actions with similar motion trajectories can be more accurately differentiated. We explicitly introduce high-level semantics of nodes in the network (node types and frame indexes) to improve the feature representation of nodes. To significantly improve the spatial correlation essential for accurate data interpretation, we present a sophisticated spatial Data-driven excitation (SDE) module. Moreover, integrating a streamlined Multi-scale Temporal Convolution (MS-TCN) network serves to significantly strengthen the representation of temporal features. Our method has been meticulously validated against the NTU-RGB D 60 and NTU-RGB D 120 datasets, and the findings confirm its superiority over existing state-of-the-art methods in terms of accuracy and robustness.