Action recognition and localization in a video are challenging tasks in video analysis, requiring detecting and localizing actions within video sequences. Recent research has increasingly focused on enhancing the modeling of long-term temporal context. To address the said task in this paper, we have proposed a novel project and pool architecture. The proposed architecture comprises of three modules. In the initial module, we proposed LSTMProjector, which is a two-layer long-short-term memory module that projects spatial and temporal features from extracted videos in feature space. It efficiently handles input features by leveraging both local and global context information. The first LSTM layer processes each feature channel independently to capture local spatial dependencies, while the second layer captures global temporal dependencies across the entire sequence. In the second module, we devise a latent space projection technique to project the extracted features into a latent space using a one-dimensional convolutional layer to match the dimensions of spatial and temporal features. In the final module, a temporal pooling module is designed, which is a parameter-free max-pooling block and operates on local regions. It enhances the efficiency of the action localization model by selectively extracting the most crucial information from neighbouring and local clip embedding. We have demonstrated the effectiveness of the proposed scheme, using mean average precision (mAP) over different thresholds of Intersection over Union (IoU) on the “THUMOS 14”, “Epic-Kitchens”, and “MultiTHUMOS” datasets. The proposed technique achieves a mean average precision (mAP) of 67.38% on THUMOS 14, 24.71% on Epic-Kitchens verb, 23.30% on Epic-Kitchens noun, and 29.91% on MultiTHUMOS datasets. Additionally, we compared the performance of the proposed techniques with those of the twenty-eight state-of-the-art (SOTA) techniques on different benchmark databases, confirming the superiority of our proposed scheme.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Project and Pool: An Action Localization Network for Localizing Actions in Untrimmed Videos

  • Himanshu Singh,
  • Avijit Dey,
  • Badri Narayan Subudhi,
  • Vinit Jakhetiya

摘要

Action recognition and localization in a video are challenging tasks in video analysis, requiring detecting and localizing actions within video sequences. Recent research has increasingly focused on enhancing the modeling of long-term temporal context. To address the said task in this paper, we have proposed a novel project and pool architecture. The proposed architecture comprises of three modules. In the initial module, we proposed LSTMProjector, which is a two-layer long-short-term memory module that projects spatial and temporal features from extracted videos in feature space. It efficiently handles input features by leveraging both local and global context information. The first LSTM layer processes each feature channel independently to capture local spatial dependencies, while the second layer captures global temporal dependencies across the entire sequence. In the second module, we devise a latent space projection technique to project the extracted features into a latent space using a one-dimensional convolutional layer to match the dimensions of spatial and temporal features. In the final module, a temporal pooling module is designed, which is a parameter-free max-pooling block and operates on local regions. It enhances the efficiency of the action localization model by selectively extracting the most crucial information from neighbouring and local clip embedding. We have demonstrated the effectiveness of the proposed scheme, using mean average precision (mAP) over different thresholds of Intersection over Union (IoU) on the “THUMOS 14”, “Epic-Kitchens”, and “MultiTHUMOS” datasets. The proposed technique achieves a mean average precision (mAP) of 67.38% on THUMOS 14, 24.71% on Epic-Kitchens verb, 23.30% on Epic-Kitchens noun, and 29.91% on MultiTHUMOS datasets. Additionally, we compared the performance of the proposed techniques with those of the twenty-eight state-of-the-art (SOTA) techniques on different benchmark databases, confirming the superiority of our proposed scheme.