Fustar: Divide and Conquer Query in Video Retrieval System
摘要
Video retrieval is a crucial task as the volume of multimedia data grows rapidly, requiring advanced systems to handle complex and diverse queries. In this paper, we present FUSTAR, a video retrieval system developed for the Ho Chi Minh AI Challenge 2024. FUSTAR leverages features from the CLIP-based model, combined with OCR, ASR, and object-based detection, to support robust text-based search. We also implemented an improved version of clustering algorithm to optimize the number of keyframes extracted. Besides, We propose two key innovations: fused search, which combines multiple features to improve accuracy and handle long queries, and temporal search, designed to process time-based video queries. Additionally, we integrate human feedback re-ranking and query refinement using large language models (LLMs) to enhance result relevance. FUSTAR’s user-friendly interface ensures accessibility for non-technical users, offering a powerful yet simple tool for effective video retrieval.