Human Action Recognition (HAR) plays an important role in various applications such as video surveillance, human-computer interaction, and healthcare. In recent years, there has been increasing interest in using skeleton-based representations for HAR due to their robustness to changes in appearance and viewpoints. However, skeleton-based approaches suffer from the class similarity problem due to the high intraclass variability and low interclass separability, because of the inherent structure of the human skeleton. Graph Convolutional Networks (GCNs) have reached remarkable results, modeling spatio-temporal relationships of a skeleton sequence. Motivated by the class similarity problem and the lack of a GCN optimization tool for skeleton-based HAR context, we propose a simple gradient-based re-training pipeline applicable to any state-of-the-art GCN model that constrains low-confidence predictions, guiding the model towards the ground-truth categories. By focusing on the relevant frames and nodes of each category, certain incorrect patterns yielded by these low-confidence samples are ignored, leading to a notable optimization of the model. Experimental results demonstrate the flexibility and effectiveness of the proposed method, improving up to 3% the accuracy of mainstream GCN recognizers.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Spatio-Temporal Deviation Calibration for Skeleton-Based Human Action Recognition

  • Gerard Marcos Freixas,
  • Zunlei Feng,
  • Kelvin Ting Zuo Han,
  • Cheng Jin,
  • Jiacong Hu,
  • Jie Lei,
  • Xingjiao Wu

摘要

Human Action Recognition (HAR) plays an important role in various applications such as video surveillance, human-computer interaction, and healthcare. In recent years, there has been increasing interest in using skeleton-based representations for HAR due to their robustness to changes in appearance and viewpoints. However, skeleton-based approaches suffer from the class similarity problem due to the high intraclass variability and low interclass separability, because of the inherent structure of the human skeleton. Graph Convolutional Networks (GCNs) have reached remarkable results, modeling spatio-temporal relationships of a skeleton sequence. Motivated by the class similarity problem and the lack of a GCN optimization tool for skeleton-based HAR context, we propose a simple gradient-based re-training pipeline applicable to any state-of-the-art GCN model that constrains low-confidence predictions, guiding the model towards the ground-truth categories. By focusing on the relevant frames and nodes of each category, certain incorrect patterns yielded by these low-confidence samples are ignored, leading to a notable optimization of the model. Experimental results demonstrate the flexibility and effectiveness of the proposed method, improving up to 3% the accuracy of mainstream GCN recognizers.