TeleoWatch: Pose-Transformer-Based Advanced Action Recognition
摘要
Recognition of human actions from videos is a valuable application for building management, security systems, accident prevention, accident intervention and several other applications. This study provides a framework for joint-based action recognition for human and non-human action recognition and common use cases for advanced intelligent closed circuit television management systems. This study offers a scaling and normalization algorithm that makes the model position agnostic for the person’s position in the frame and allows for generalization for actions performed facing the camera at various angles. It also improves on previous pose-based action recognition systems (Bidirectional-LSTM) by using a transformer encoder architecture for the recognition task. The transformer encoder architecture uses the joint locations output from pose estimation computer vision models to train the model to recognise the common joint trajectories that constitute human actions. The proposed transformer encoder model provides better performance than bidirectional-LSTM models in both speed and accuracy and provides better generalisation to novel tasks on benchmark action recognition datasets such as the KTH dataset and the UR-Fall dataset.