Video retrieval, the process of locating specific video content within large datasets, presents a significant challenge in the era of digital multimedia. In response to this, as part of the Ho Chi Minh City AI Challenge 2024, this paper presents an advanced multi-modalities and multi-stages video retrieval framework namely MMMSVR, enhanced with the ability to answer supplemental questions, such as automatic counting of objects. The proposed method leverages vision-language models, combining CLIP ViT-H/14, BLIP2 and BEiT-3 for feature encoding and implements a re-ranking mechanism based on a weighting system. Furthermore, a wide range of query modalities such as Optical Character Recognition (OCR), Object Detection, and Automatic Speech Recognition are integrated to refine and improve the retrieval process, supporting both text and image-based queries for efficient retrieval based on multiple attributes. The framework also includes an image-based query feature, enriching the model’s versatility and improving retrieval accuracy. The proposed approach demonstrates significant performance improvements and offers a robust, flexible solution for video search and question-answering tasks.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MMMSVR: An Advanced Video Retrieval and Question Answering System

  • Viet Hang Duong,
  • Tong Dang Khoa Huynh,
  • Minh Quan Tran,
  • Nguyen Khang Nguyen,
  • Canh Nhat Le,
  • Thanh Hung Nguyen

摘要

Video retrieval, the process of locating specific video content within large datasets, presents a significant challenge in the era of digital multimedia. In response to this, as part of the Ho Chi Minh City AI Challenge 2024, this paper presents an advanced multi-modalities and multi-stages video retrieval framework namely MMMSVR, enhanced with the ability to answer supplemental questions, such as automatic counting of objects. The proposed method leverages vision-language models, combining CLIP ViT-H/14, BLIP2 and BEiT-3 for feature encoding and implements a re-ranking mechanism based on a weighting system. Furthermore, a wide range of query modalities such as Optical Character Recognition (OCR), Object Detection, and Automatic Speech Recognition are integrated to refine and improve the retrieval process, supporting both text and image-based queries for efficient retrieval based on multiple attributes. The framework also includes an image-based query feature, enriching the model’s versatility and improving retrieval accuracy. The proposed approach demonstrates significant performance improvements and offers a robust, flexible solution for video search and question-answering tasks.