TFF-temporal fusion framework for advancing video retrieval through long-range dependencies and multi-modal intent
摘要
Efficiently navigating the surge in video content, our paper introduces the Temporal Fusion Framework (TFF) a pioneering approach to video content retrieval and moment detection via natural language queries. TFF utilizes a novel temporal context modeling strategy, adeptly capturing global context with long-range dependencies across diverse time scales. The integration of a specialized encoder-decoder model with multi-modal intent facilitates the extraction of intricate temporal patterns and user intent understanding. Tackling dynamic background changes, a common challenge in moment retrieval, TFF's temporal context modeling ensures accurate content retrieval. Through comprehensive comparative studies on public datasets, employing Recall (R@1, R@5), and mean average precision (mAP) as metrics, our experimental results showcase TFF's superior performance over existing methods. The incorporation of multi-modal intent and innovative temporal context modeling significantly enhances the framework's capability to identify relevant moments, marking a transformative step in video content interaction.