Spatio-Temporal Deviation Calibration for Skeleton-Based Human Action Recognition
摘要
Human Action Recognition (HAR) plays an important role in various applications such as video surveillance, human-computer interaction, and healthcare. In recent years, there has been increasing interest in using skeleton-based representations for HAR due to their robustness to changes in appearance and viewpoints. However, skeleton-based approaches suffer from the class similarity problem due to the high intraclass variability and low interclass separability, because of the inherent structure of the human skeleton. Graph Convolutional Networks (GCNs) have reached remarkable results, modeling spatio-temporal relationships of a skeleton sequence. Motivated by the class similarity problem and the lack of a GCN optimization tool for skeleton-based HAR context, we propose a simple gradient-based re-training pipeline applicable to any state-of-the-art GCN model that constrains low-confidence predictions, guiding the model towards the ground-truth categories. By focusing on the relevant frames and nodes of each category, certain incorrect patterns yielded by these low-confidence samples are ignored, leading to a notable optimization of the model. Experimental results demonstrate the flexibility and effectiveness of the proposed method, improving up to 3% the accuracy of mainstream GCN recognizers.