Partially Relevant Video Retrieval (PRVR) is a task dedicated to retrieving untrimmed videos that are partially relevant to a given textual query. In the large-scale video retrieval task, efficiency is just as crucial as accuracy. However, existing studies in PRVR typically utilize a combination of frame-level and clip-level features, which introduces significant feature redundancy, leading to increased space and time complexity at the retrieval stage. In this paper, we present a novel approach, Efficient Partially Relevant Video Retrieval with disentangled learning (EPRVR) to improve video feature representation efficiency and reduce the space and time complexity issue at the retrieval stage, particularly for large-scale video datasets. EPRVR aims to succinctly describe untrimmed videos using a minimal set of video-level features, consisting of two key components: the Partially Relevant Disentangled Network (PRDN) and Disentangled-Match Learning (DML). The PRDN is an end-to-end network that encodes videos to extract cross-modal semantic disentangled features as the video-level embedding. The DML employs a training approach that integrates a Text-Video matching loss for cross-modal semantic feature alignment and an Intra-Video Disentangle loss (IVD) to capture critical disentangled information during video encoding. Extensive experiments demonstrate that our method outperforms other video-level approaches, closely approaching the SOTA performance of multi-level (frame/clip) methods, and exhibits impressive efficiency with low memory usage and retrieval time at the retrieval stage, ensuring its practicality for large-scale and real-time video retrieval.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

EPRVR: Efficient Partially Relevant Video Retrieval with Disentangled Video Representation Learning

  • Zhaoxing Li,
  • Zhiming Zhang,
  • Dongqin Liu,
  • Yesheng Chai,
  • Jiao Dai,
  • Bibo Tu,
  • Jizhong Han

摘要

Partially Relevant Video Retrieval (PRVR) is a task dedicated to retrieving untrimmed videos that are partially relevant to a given textual query. In the large-scale video retrieval task, efficiency is just as crucial as accuracy. However, existing studies in PRVR typically utilize a combination of frame-level and clip-level features, which introduces significant feature redundancy, leading to increased space and time complexity at the retrieval stage. In this paper, we present a novel approach, Efficient Partially Relevant Video Retrieval with disentangled learning (EPRVR) to improve video feature representation efficiency and reduce the space and time complexity issue at the retrieval stage, particularly for large-scale video datasets. EPRVR aims to succinctly describe untrimmed videos using a minimal set of video-level features, consisting of two key components: the Partially Relevant Disentangled Network (PRDN) and Disentangled-Match Learning (DML). The PRDN is an end-to-end network that encodes videos to extract cross-modal semantic disentangled features as the video-level embedding. The DML employs a training approach that integrates a Text-Video matching loss for cross-modal semantic feature alignment and an Intra-Video Disentangle loss (IVD) to capture critical disentangled information during video encoding. Extensive experiments demonstrate that our method outperforms other video-level approaches, closely approaching the SOTA performance of multi-level (frame/clip) methods, and exhibits impressive efficiency with low memory usage and retrieval time at the retrieval stage, ensuring its practicality for large-scale and real-time video retrieval.