The classical Video Grounding setting has received abundant study, and state-of-the-art models are now capable of achieving remarkable performance. Nevertheless, there remain several significant issues to be explored. In this chapter, we aim to introduce these directions, including out-of-distribution Video Grounding, explainable Video Grounding, and the application of vision-language pre-training to Video Grounding. Preliminary investigations have been conducted on some of these topics, so we will summarize the early works and point out possible future directions simultaneously.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Future Research Directions

  • Xin Wang,
  • Xiaohan Lan,
  • Wenwu Zhu

摘要

The classical Video Grounding setting has received abundant study, and state-of-the-art models are now capable of achieving remarkable performance. Nevertheless, there remain several significant issues to be explored. In this chapter, we aim to introduce these directions, including out-of-distribution Video Grounding, explainable Video Grounding, and the application of vision-language pre-training to Video Grounding. Preliminary investigations have been conducted on some of these topics, so we will summarize the early works and point out possible future directions simultaneously.