MAVEN: Video Retrieval System Using a Multi-agent Visual Exploration Network
摘要
Effective video retrieval systems are essential as video data grows across various fields. Traditionally, these systems rely on OCR, object detection, color extraction, and audio analysis. Current approaches like CLIP bridge the text-image embeddings gap for search, but they often lack contextual depth for complex, multi-frame searches. We propose a solution that integrates traditional methods and CLIP with advanced language models and prompting techniques for image captioning, extracting rich information from individual frames. Our system includes an Agent that automates searches, classifies queries, generates prompts, and verifies results, improving search accuracy while reducing user effort. Additional features like temporal search, video previews, and frame filtering further enhance the user experience. This comprehensive approach provides a powerful toolkit for achieving more accurate and efficient video search results, addressing the growing complexity of video data retrieval across various domains.