The rapid growth of video content on social media platforms, surveillance systems, and educational repositories has made automatic video summarization essential for efficient content consumption and analysis. In this work, we propose a lightweight Transformer-based encoder–decoder model tailored for video summarization. Our approach combines the Transformer’s ability to capture long-range temporal dependencies with architectural optimizations that significantly reduce computational overhead. The model features a compact encoder–decoder structure, a learnable start token, and causal self-attention in the decoder to generate frame-level importance scores autoregressively. Despite having only 2.9 M parameters, the model achieves a state-of-the-art F1-score of approximately 83.22% on the TVSum dataset. It outperforms existing methods while maintaining fast inference, making it well-suited for real-time applications. We evaluate our model under three experimental modes: canonical, augmented, and transfer, and observe consistent performance across all settings. Hyperparameters are optimized using Bayesian search via Optuna to ensure stability and generalizability. Our findings show that high-quality summarization can be achieved without large-scale architectures. The compactness, speed, and accuracy of our model position it as a practical and robust solution for modern video summarization needs.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Lightweight Transformer-Based Encoder-Decoder Model for Video Summarization

  • Saadman Sakib,
  • Rajesh Palit,
  • Dipankar Das,
  • Tanjim Mahmud,
  • Kaushik Deb

摘要

The rapid growth of video content on social media platforms, surveillance systems, and educational repositories has made automatic video summarization essential for efficient content consumption and analysis. In this work, we propose a lightweight Transformer-based encoder–decoder model tailored for video summarization. Our approach combines the Transformer’s ability to capture long-range temporal dependencies with architectural optimizations that significantly reduce computational overhead. The model features a compact encoder–decoder structure, a learnable start token, and causal self-attention in the decoder to generate frame-level importance scores autoregressively. Despite having only 2.9 M parameters, the model achieves a state-of-the-art F1-score of approximately 83.22% on the TVSum dataset. It outperforms existing methods while maintaining fast inference, making it well-suited for real-time applications. We evaluate our model under three experimental modes: canonical, augmented, and transfer, and observe consistent performance across all settings. Hyperparameters are optimized using Bayesian search via Optuna to ensure stability and generalizability. Our findings show that high-quality summarization can be achieved without large-scale architectures. The compactness, speed, and accuracy of our model position it as a practical and robust solution for modern video summarization needs.