Transforming Video Search: Leveraging Multimodal Techniques and LLMs for Optimal Retrieval
摘要
The rapid growth of online video content has created an urgent need for efficient and accurate event-based video retrieval systems. Existing techniques, such as image-text retrieval, audio analysis, and text-based searches, frequently fail to cope with complex video data and extract useful information from multiple modalities. This paper presents the Multimodal Mapping and Retrieval System (MMRS-LMF), which uses Large Language Models and Multi-Stage Fusion. This novel system improves video retrieval by combining multimodal content (video, audio, and text) into a single, text-based format. It improves retrieval precision and recall by utilizing advanced text embedding techniques and multimodal fusion. The experimental results show significant improvements in retrieval accuracy across a variety of video datasets, demonstrating the system’s ability to meet the needs of modern event-based video search applications.