<p>In contrast to two-dimensional skeletal data, which exhibits inherent susceptibility to environmental interference such as complex backgrounds and illumination variations, three-dimensional skeletal data demonstrates superior robustness against these confounding factors. This enhanced spatial representation capability enables more accurate feature extraction, thereby offering a more reliable framework for human action recognition tasks. However, feeding long videos into recognition models can significantly increase computational costs and potentially degrade accuracy due to redundant information. To address this, we propose a novel method based on meta-process-driven feature learning from 3D skeleton data. Our approach divides videos into meta-processes through clustering based on motion features of core joint points. Each meta-process is characterized using special Euclidean groups, and their parameters are derived using a Transformer network. By compressing data representation to eliminate redundancy and training a classifier with fused internal and external meta-process features, our method achieves efficient and accurate action recognition. Experimental results on NTU-RGB+D60, NTU-RGB+D120 and a mixed dataset demonstrate the effectiveness and advantages of our approach compared to 18 state-of-the-art algorithms.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Meta-process-driven 3D skeleton feature learning for enhanced human action recognition

  • Hao Chen,
  • Chenxi Wang,
  • Jiawei Duan,
  • Qing Yang

摘要

In contrast to two-dimensional skeletal data, which exhibits inherent susceptibility to environmental interference such as complex backgrounds and illumination variations, three-dimensional skeletal data demonstrates superior robustness against these confounding factors. This enhanced spatial representation capability enables more accurate feature extraction, thereby offering a more reliable framework for human action recognition tasks. However, feeding long videos into recognition models can significantly increase computational costs and potentially degrade accuracy due to redundant information. To address this, we propose a novel method based on meta-process-driven feature learning from 3D skeleton data. Our approach divides videos into meta-processes through clustering based on motion features of core joint points. Each meta-process is characterized using special Euclidean groups, and their parameters are derived using a Transformer network. By compressing data representation to eliminate redundancy and training a classifier with fused internal and external meta-process features, our method achieves efficient and accurate action recognition. Experimental results on NTU-RGB+D60, NTU-RGB+D120 and a mixed dataset demonstrate the effectiveness and advantages of our approach compared to 18 state-of-the-art algorithms.