This paper presents ViFi - a Video Finding System for Video Browser Showdown 2025. Our retrieval system is mainly based on the SigLIP, a most recent and robust visual-textual embedding model, and significantly outperforms CLIP in text-to-image retrieval on the MSCOCO data. Instead of using the given extracted keyframe from the dataset like other previous approaches, we found that the given keyframe might have missed some of the moments in the textual query. Thus, our system operates using keyframes derived from our custom keyframe extraction stage. While neighboring keyframes might share similar characteristics, such as objects’ appearance, we also try to visualize the result in a relevant video segment extracted by TransNetv2. Moreover, our system supports temporal search with now-and-then textual input, allowing users to search for time-related information. Finally, we utilized recent tools to extract metadata from the keyframes, including detected objects, the count of these objects, and more. We then combined this information with user input to refine and filter the search results. Finally, we have designed a user-friendly interface offering both simple and advanced modes to enable faster examinations and assist novice users.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ViFi: A Video Finding System at Video Browser Showdown 2025

  • Khanh-An C. Quan,
  • Qui Ngoc Nguyen,
  • Minh-Triet Tran

摘要

This paper presents ViFi - a Video Finding System for Video Browser Showdown 2025. Our retrieval system is mainly based on the SigLIP, a most recent and robust visual-textual embedding model, and significantly outperforms CLIP in text-to-image retrieval on the MSCOCO data. Instead of using the given extracted keyframe from the dataset like other previous approaches, we found that the given keyframe might have missed some of the moments in the textual query. Thus, our system operates using keyframes derived from our custom keyframe extraction stage. While neighboring keyframes might share similar characteristics, such as objects’ appearance, we also try to visualize the result in a relevant video segment extracted by TransNetv2. Moreover, our system supports temporal search with now-and-then textual input, allowing users to search for time-related information. Finally, we utilized recent tools to extract metadata from the keyframes, including detected objects, the count of these objects, and more. We then combined this information with user input to refine and filter the search results. Finally, we have designed a user-friendly interface offering both simple and advanced modes to enable faster examinations and assist novice users.