Recognition of human actions from videos is a valuable application for building management, security systems, accident prevention, accident intervention and several other applications. This study provides a framework for joint-based action recognition for human and non-human action recognition and common use cases for advanced intelligent closed circuit television management systems. This study offers a scaling and normalization algorithm that makes the model position agnostic for the person’s position in the frame and allows for generalization for actions performed facing the camera at various angles. It also improves on previous pose-based action recognition systems (Bidirectional-LSTM) by using a transformer encoder architecture for the recognition task. The transformer encoder architecture uses the joint locations output from pose estimation computer vision models to train the model to recognise the common joint trajectories that constitute human actions. The proposed transformer encoder model provides better performance than bidirectional-LSTM models in both speed and accuracy and provides better generalisation to novel tasks on benchmark action recognition datasets such as the KTH dataset and the UR-Fall dataset.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

TeleoWatch: Pose-Transformer-Based Advanced Action Recognition

  • Hanno Jacobs,
  • Thambo Nyathi

摘要

Recognition of human actions from videos is a valuable application for building management, security systems, accident prevention, accident intervention and several other applications. This study provides a framework for joint-based action recognition for human and non-human action recognition and common use cases for advanced intelligent closed circuit television management systems. This study offers a scaling and normalization algorithm that makes the model position agnostic for the person’s position in the frame and allows for generalization for actions performed facing the camera at various angles. It also improves on previous pose-based action recognition systems (Bidirectional-LSTM) by using a transformer encoder architecture for the recognition task. The transformer encoder architecture uses the joint locations output from pose estimation computer vision models to train the model to recognise the common joint trajectories that constitute human actions. The proposed transformer encoder model provides better performance than bidirectional-LSTM models in both speed and accuracy and provides better generalisation to novel tasks on benchmark action recognition datasets such as the KTH dataset and the UR-Fall dataset.