LLM-Powered Video Search: A Comprehensive Multimedia Retrieval System
摘要
Image and video search has become an important problem amid the rapid growth of image and video-sharing platforms. The explosion of video content on the Internet has created an urgent demand for effective content management and retrieval. Traditional approaches, such as text-based, image, object, audio, and color searches, have achieved some success but often lack consistency and fail to meet user needs comprehensively. To address this issue, we propose an artificial intelligence (AI) system that integrates multiple existing search methods with a Combined Ranking Score (CRS) algorithm. CRS balances the ranking and scores from different approaches, optimizing query results. The system also uses a Retrieval-Augmented Generation (RAG) architecture combined with large language models (LLMs) to enhance contextual understanding and deliver more accurate results. Users can interact directly with our AI system to easily and efficiently search for desired moments in videos. This approach promises to significantly improve the user experience in multimedia content search and retrieval.