<p>Video Transformer has become the standard architecture for video recognition due to its powerful spatio-temporal modeling capability. However, its high computational cost makes it difficult to apply in real-world scenarios. In this paper, we design a lightweight module called Midframe-Centric Token Merging (MCTM). It can be embedded in some video transformers to adaptively merge some unimportant and similar tokens during the inference process. Specifically, MCTM can be divided into two progressive sub-modules. In the first sub-module, we divide tokens into two parts based on their importance, called important tokens and unimportant tokens. In the second sub-module, for unimportant tokens, we design a token merging strategy to merge some similar tokens. We performed several sets of experiments to validate the effectiveness of MCTM. For example, we embedded MCTM in Joint Video Transformer and ran experiments on Kinetics-400 dataset. It reduces the computational cost by 37% while maintaining similar accuracy.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Midframe-centric token merging for efficient video transformer

  • Qian Zhang,
  • Zuosui Yang,
  • Mingwen Shao,
  • Hong Liang

摘要

Video Transformer has become the standard architecture for video recognition due to its powerful spatio-temporal modeling capability. However, its high computational cost makes it difficult to apply in real-world scenarios. In this paper, we design a lightweight module called Midframe-Centric Token Merging (MCTM). It can be embedded in some video transformers to adaptively merge some unimportant and similar tokens during the inference process. Specifically, MCTM can be divided into two progressive sub-modules. In the first sub-module, we divide tokens into two parts based on their importance, called important tokens and unimportant tokens. In the second sub-module, for unimportant tokens, we design a token merging strategy to merge some similar tokens. We performed several sets of experiments to validate the effectiveness of MCTM. For example, we embedded MCTM in Joint Video Transformer and ran experiments on Kinetics-400 dataset. It reduces the computational cost by 37% while maintaining similar accuracy.