InVideo Search: Scene Description Clustering and Integrating Image and Audio Captioning for Enhanced Video Search
摘要
In the digital era, the vast amount of long videos highlights the need for an advanced system to efficiently find specific segments based on text descriptions. The proposed solution aims to develop a video content retrieval system by utilizing established tools in video analysis, image/video captioning, and information retrieval. The work is organized into three modules: Keyframe extraction, Captioning, and Query-based searching. Our primary focus is on effectively identifying key elements within these videos, achieving a semantic understanding of these videos in the form of captions, and establishing precise matches of user queries. In short, users input descriptive text prompts, and the application, equipped with pre-trained models, promptly and accurately identifies and retrieves the video segments. This application carries significant potential across various domains, from enabling efficient scene location in movies to helping extract knowledge from educational lectures. Furthermore, this research fundamentally departs from conventional video search methods that heavily rely on metadata or manually assigned tags, which often fall short of capturing the nuances of video content. This work also focuses on a multimodal captioning strategy to integrate more semantic content. Clustering of captions was also incorporated to improve search performance in the case of long videos. This approach allows us to deliver a robust and efficient solution, taking advantage of the extensive knowledge and research underlying these established technologies. Beyond the immediate applications, this research has broader applications in fields like natural language understanding, multimedia analysis, and information retrieval.