Temporal Moment Localization (TML) identifies specific temporal intervals in untrimmed videos based on a sentence query. Traditional methods using 2D temporal maps face limitations due to rigid boundaries and GPU constraints. We propose a Boundary Matching and Refinement Network (BMRN) that dynamically adjusts moment proposals with predicted center and length offsets for precise localization. BMRN integrates boundary matching and refinement maps with a length-aware cross-modal interactive proposal feature map. Enhanced with Cross-Modal Contrastive Learning (CCL), BMRN-CCL reduces the impact of visually and semantically similar negative samples. Extensive ablation studies and benchmarks on Charades-STA and ActivityNet Captions datasets demonstrate the superior performance of BMRN and BMRN-CCL, surpassing state-of-the-art methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Boundary Matching and Refinement Network with Cross-Modal Contrastive Learning for Temporal Moment Localization

  • Jinyoung Moon,
  • Muah Seol,
  • Jonghee Kim

摘要

Temporal Moment Localization (TML) identifies specific temporal intervals in untrimmed videos based on a sentence query. Traditional methods using 2D temporal maps face limitations due to rigid boundaries and GPU constraints. We propose a Boundary Matching and Refinement Network (BMRN) that dynamically adjusts moment proposals with predicted center and length offsets for precise localization. BMRN integrates boundary matching and refinement maps with a length-aware cross-modal interactive proposal feature map. Enhanced with Cross-Modal Contrastive Learning (CCL), BMRN-CCL reduces the impact of visually and semantically similar negative samples. Extensive ablation studies and benchmarks on Charades-STA and ActivityNet Captions datasets demonstrate the superior performance of BMRN and BMRN-CCL, surpassing state-of-the-art methods.