<p>Human action recognition (HAR) is an emerging research area with a diverse range of applications in surveillance, healthcare, and other fields. However, HAR in surveillance systems presents challenges due to the need for real-time processing of video data in dynamic environments. Capturing both spatial and temporal dependencies in surveillance footage is difficult, especially when dealing with cluttered backgrounds, diverse viewpoints, and partial occlusions. These challenges arise from the computational limitations of many surveillance systems, which make the use of computationally intensive models difficult for efficient and timely activity recognition. To overcome these challenges, we propose a novel and efficient hybrid framework for robust HAR, specifically designed for resource-constrained devices. EfficientNet-B0 up to the block 5 layer with the set of salient contextual features and dimensions of 7x7x1280 is leveraged to extract spatial features from individual frames. By extracting features up to the blocks 5 layers, the model efficiently balances the trade-off between performance and computational cost while still capturing rich spatial representations. In order to understand long-range temporal dependencies, the extracted spatial features vector is then sent to a vision transformer (ViT), which analyzes the sequential frame features in the second stage. This approach enables efficient modeling of temporal relationships across frames. A focal loss function is applied to handle class imbalance, further enhancing the model’s robustness and performance. The proposed framework is evaluated on three challenging HAR datasets-UCF50, YouTube Action, and HMDB51-achieving competitive accuracy and outperforming existing state-of-the-art methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

E-harnet: an efficient hybrid transformer network for human activity recognition

  • Asif Iqbal,
  • Muhammad Arslan Rauf,
  • Salim,
  • M. D. Shakib Mahamud,
  • Mian Muhammad Yasir Khalil,
  • Zhen Qin

摘要

Human action recognition (HAR) is an emerging research area with a diverse range of applications in surveillance, healthcare, and other fields. However, HAR in surveillance systems presents challenges due to the need for real-time processing of video data in dynamic environments. Capturing both spatial and temporal dependencies in surveillance footage is difficult, especially when dealing with cluttered backgrounds, diverse viewpoints, and partial occlusions. These challenges arise from the computational limitations of many surveillance systems, which make the use of computationally intensive models difficult for efficient and timely activity recognition. To overcome these challenges, we propose a novel and efficient hybrid framework for robust HAR, specifically designed for resource-constrained devices. EfficientNet-B0 up to the block 5 layer with the set of salient contextual features and dimensions of 7x7x1280 is leveraged to extract spatial features from individual frames. By extracting features up to the blocks 5 layers, the model efficiently balances the trade-off between performance and computational cost while still capturing rich spatial representations. In order to understand long-range temporal dependencies, the extracted spatial features vector is then sent to a vision transformer (ViT), which analyzes the sequential frame features in the second stage. This approach enables efficient modeling of temporal relationships across frames. A focal loss function is applied to handle class imbalance, further enhancing the model’s robustness and performance. The proposed framework is evaluated on three challenging HAR datasets-UCF50, YouTube Action, and HMDB51-achieving competitive accuracy and outperforming existing state-of-the-art methods.