Previous approaches to the task of video question answering (VideoQA) rely on either frame-level representations or region-level representations to encode visual contents of a video. Regarding the former stream of approaches, one strategy is to utilize all the frames selected according to a fixed time interval and the other strategy is to select a single frame to stand for the whole video. In this paper, we repel these two extremes and focus on the question of whether dynamically selecting frames can achieve better performance. To this end, we propose a novel weakly-supervised approach to dynamic frame selection. In addition, to make better use of selected frames, we utilize a multi-instance learning method to learn to predict answers based on frame-question pairs rather than the holistic representation of all the selected frames. Extensive experiments on benchmarks show that the proposed approach achieves the best or competitive performance, which is very promising especially when we take its relative simplicity into consideration.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DFS-QA: Dynamic Frame Selection for Better Video Question Answering

  • Zhibo Ren,
  • Baoyu Hou,
  • Huizhen Wang,
  • Muhua Zhu,
  • Tong Xiao,
  • Jingbo Zhu

摘要

Previous approaches to the task of video question answering (VideoQA) rely on either frame-level representations or region-level representations to encode visual contents of a video. Regarding the former stream of approaches, one strategy is to utilize all the frames selected according to a fixed time interval and the other strategy is to select a single frame to stand for the whole video. In this paper, we repel these two extremes and focus on the question of whether dynamically selecting frames can achieve better performance. To this end, we propose a novel weakly-supervised approach to dynamic frame selection. In addition, to make better use of selected frames, we utilize a multi-instance learning method to learn to predict answers based on frame-question pairs rather than the holistic representation of all the selected frames. Extensive experiments on benchmarks show that the proposed approach achieves the best or competitive performance, which is very promising especially when we take its relative simplicity into consideration.