<p>In this paper, we propose a numerical spatiotemporal approach for video summarization. The current solutions leverage deep learning techniques to tackle this task. However, existing methods do not employ the video shots’ length data in their networks. We first introduce a new ground truth labelling for video summarization. This ground truth tackles multiple users’ annotations and is inclusive of video shot length information. We then propose a novel network architecture Numerical SpatioTemporal Fusion Net, NSTFNet. The proposed architecture leverages the temporal modelling ability of the self-attention mechanism to model video’s vision data to be fused with structured numerical features of video shots. We evaluate our results on TVSum and SumMe datasets. Experimental results show that the proposed model outperforms state-of-the-art performance on the SumMe dataset with 51.29% and 54.42% f-scores on the canonical and augmented configurations, demonstrating the effectiveness of our proposed model compared to state-of-the-art methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Numerical and spatiotemporal features fusion for video summarization

  • Mohamed Aboelenien,
  • Mohammed A.-M. Salem

摘要

In this paper, we propose a numerical spatiotemporal approach for video summarization. The current solutions leverage deep learning techniques to tackle this task. However, existing methods do not employ the video shots’ length data in their networks. We first introduce a new ground truth labelling for video summarization. This ground truth tackles multiple users’ annotations and is inclusive of video shot length information. We then propose a novel network architecture Numerical SpatioTemporal Fusion Net, NSTFNet. The proposed architecture leverages the temporal modelling ability of the self-attention mechanism to model video’s vision data to be fused with structured numerical features of video shots. We evaluate our results on TVSum and SumMe datasets. Experimental results show that the proposed model outperforms state-of-the-art performance on the SumMe dataset with 51.29% and 54.42% f-scores on the canonical and augmented configurations, demonstrating the effectiveness of our proposed model compared to state-of-the-art methods.