VClipper: Moment Retrieval in Video Streams Using Zero-Shot and Context-Aware Foundational Models
摘要
Widespread access to digital cameras has led to a significant increase of photo and video editing tools. Furthermore, sharing such content has led to the popularization of dozens of social media websites and apps. In this research we explore the extent to which a well-known foundational model, such as CLIP -Contrastive Language-Image Pre-Training- model, a pretrained model with semantic and scene recognition capabilities, can be used, without further training, as a moment retrieval searcher in video recordings. We propose two novel methods aimed at moment retrieval tasks in audiovisual data, namely VClipper-frame and VClipper-scene, and perform an empirical analysis on the effects of post-processing on the similarity vectors obtained from CLIP encoders. Our methods for zero-shot moment retrieval using CLIP scored better than the current state-of-the-art. Specifically, VClipper-frame reaches \(R@1_{0.5}=57.4\% \) and \(mAP_{0.5}=51.6\%\) . Compared to previous work, that reached values \(R@1_{0.5}=42.1\% \) and \(mAP_{0.5}=43.0\%\) , our approach represents an improvement of 15.3 points and 8.6 points, respectively.