Temporal localization of moments in videos, driven by natural language queries, is essential for applications like video search and summarization. Transformer-based models such as Moment Detr, CG Detr, and MH Detr have advanced performance on datasets like QV-Highlights and Charades-STA. However, their reliance on generic feature extractors often overlooks dataset-specific nuances. We propose a dataset-driven approach that employs tailored feature extractors: I3D for QV-Highlights’ dynamic sports highlights and ResNet-50 with RPN for Charades-STA’s object-centric activities. Integrated into a unified transformer architecture, our method achieves a 1–2% improvement in Recall@1 and mIoU across all datasets compared to existing transformer-based models that rely on generic feature extractors across all datasets. Our contributions include: specialized feature extractors capturing dataset-specific patterns, a scalable transformer-based framework and new performance benchmarks validated through extensive experiments. These findings emphasize the critical role of dataset-adaptive feature extraction in enhancing moment detection at scale.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Moment Detection at Scale: Dataset-Driven Techniques for Temporal Localization

  • Md. Misbah Khan,
  • Aritra Islam Saswato,
  • Sabbir Hossain,
  • Towfiqur Rahman Toki,
  • Md. Towsif Abir,
  • Shafin Rahman

摘要

Temporal localization of moments in videos, driven by natural language queries, is essential for applications like video search and summarization. Transformer-based models such as Moment Detr, CG Detr, and MH Detr have advanced performance on datasets like QV-Highlights and Charades-STA. However, their reliance on generic feature extractors often overlooks dataset-specific nuances. We propose a dataset-driven approach that employs tailored feature extractors: I3D for QV-Highlights’ dynamic sports highlights and ResNet-50 with RPN for Charades-STA’s object-centric activities. Integrated into a unified transformer architecture, our method achieves a 1–2% improvement in Recall@1 and mIoU across all datasets compared to existing transformer-based models that rely on generic feature extractors across all datasets. Our contributions include: specialized feature extractors capturing dataset-specific patterns, a scalable transformer-based framework and new performance benchmarks validated through extensive experiments. These findings emphasize the critical role of dataset-adaptive feature extraction in enhancing moment detection at scale.