Video event retrieval task aims at retrieving video events from a large video collection that are semantically relevant to a given textual or visual query. There are several approaches have been introduced for this promising task, but overall, they can be categorized either as embedding-based techniques or concept-based techniques. Each of them has its advantages and an appropriately used context. The huge size of datasets often seen in this task is also a challenge for retrieval systems to work efficiently. This paper presents a comprehensive solution and application for video retrieval that addresses the challenges of speed and accuracy in large-scale datasets. Our approach integrates two complementary methods: an image semantic search using CLIP visual-textual embeddings together with an advanced FAISS vector retrieve index; and a concept-based search using optical character recognition (OCR) and automatic speech recognition (ASR) together with Elastic Text Search. In the first approach, we further implement four strategies that advantage and enhance the power of CLIP embeddings. Through experiments, we demonstrate that our approaches provide high accuracy and efficiency for video retrieval applications for a vast type of query. We also found that the quality of the query noticeably impacted the search results and the ability of end users to customize the search by modifying different factors also contributed to a quick and successful query.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

LameFrames: Optimizing Video Event Retrieval Through Strategic Integration and Individual Strategy Enhancement

  • Hai-Chan Nguyen,
  • Tuong-Duy Nguyen-Dang,
  • Phuong-Thuy Le-Nguyen,
  • Long Huynh,
  • Khai-Hung Do

摘要

Video event retrieval task aims at retrieving video events from a large video collection that are semantically relevant to a given textual or visual query. There are several approaches have been introduced for this promising task, but overall, they can be categorized either as embedding-based techniques or concept-based techniques. Each of them has its advantages and an appropriately used context. The huge size of datasets often seen in this task is also a challenge for retrieval systems to work efficiently. This paper presents a comprehensive solution and application for video retrieval that addresses the challenges of speed and accuracy in large-scale datasets. Our approach integrates two complementary methods: an image semantic search using CLIP visual-textual embeddings together with an advanced FAISS vector retrieve index; and a concept-based search using optical character recognition (OCR) and automatic speech recognition (ASR) together with Elastic Text Search. In the first approach, we further implement four strategies that advantage and enhance the power of CLIP embeddings. Through experiments, we demonstrate that our approaches provide high accuracy and efficiency for video retrieval applications for a vast type of query. We also found that the quality of the query noticeably impacted the search results and the ability of end users to customize the search by modifying different factors also contributed to a quick and successful query.