<p>Efficiently navigating the surge in video content, our paper introduces the Temporal Fusion Framework (TFF) a pioneering approach to video content retrieval and moment detection via natural language queries. TFF utilizes a novel temporal context modeling strategy, adeptly capturing global context with long-range dependencies across diverse time scales. The integration of a specialized encoder-decoder model with multi-modal intent facilitates the extraction of intricate temporal patterns and user intent understanding. Tackling dynamic background changes, a common challenge in moment retrieval, TFF's temporal context modeling ensures accurate content retrieval. Through comprehensive comparative studies on public datasets, employing Recall (R@1, R@5), and mean average precision (mAP) as metrics, our experimental results showcase TFF's superior performance over existing methods. The incorporation of multi-modal intent and innovative temporal context modeling significantly enhances the framework's capability to identify relevant moments, marking a transformative step in video content interaction.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

TFF-temporal fusion framework for advancing video retrieval through long-range dependencies and multi-modal intent

  • Pratibha Singh,
  • Kashvi Chakrawal,
  • Alok Kumar Singh Kushwaha

摘要

Efficiently navigating the surge in video content, our paper introduces the Temporal Fusion Framework (TFF) a pioneering approach to video content retrieval and moment detection via natural language queries. TFF utilizes a novel temporal context modeling strategy, adeptly capturing global context with long-range dependencies across diverse time scales. The integration of a specialized encoder-decoder model with multi-modal intent facilitates the extraction of intricate temporal patterns and user intent understanding. Tackling dynamic background changes, a common challenge in moment retrieval, TFF's temporal context modeling ensures accurate content retrieval. Through comprehensive comparative studies on public datasets, employing Recall (R@1, R@5), and mean average precision (mAP) as metrics, our experimental results showcase TFF's superior performance over existing methods. The incorporation of multi-modal intent and innovative temporal context modeling significantly enhances the framework's capability to identify relevant moments, marking a transformative step in video content interaction.