In response to the growing need for precise and efficient video retrieval, we present a versatile video event retrieval system that supports multiple query modalities. Our system integrates 5 modes for querying in natural languages: quick search, temporal search, hybrid text and sketch search, transcript-based search, and OCR-based search. Powered by the CLIP ViT-L model, the quick search matches user queries to relevant video segments. Temporal search combines CLIP embeddings with mathematical techniques to pinpoint specific timeframes, while the TASK-former model supports hybrid sketch-text search, enabling users to locate scenes using hand-drawn sketches and textual descriptions. Transcript-based search aids users to overcome socio-cultural barriers and OCR-based search utilizes extracted text from video keyframes. The system’s interface allows users to input queries, and browse top-ranked results in a manner that tackles clustered viewing problems seen in common systems. In addition, users can explore visually similar images, and preview short clips, improving both precision and accessibility in video content retrieval.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Comprehensive Video Event Retrieval System for Vietnamese News: Integrating CLIP ViT, TASK-Former, Transcripts, and OCR

  • Thinh-Phat Vo,
  • Quang-Thang Duong,
  • Quoc-Thang Nguyen,
  • Dang-Khoa Mai,
  • Nguyen-Khang Ly

摘要

In response to the growing need for precise and efficient video retrieval, we present a versatile video event retrieval system that supports multiple query modalities. Our system integrates 5 modes for querying in natural languages: quick search, temporal search, hybrid text and sketch search, transcript-based search, and OCR-based search. Powered by the CLIP ViT-L model, the quick search matches user queries to relevant video segments. Temporal search combines CLIP embeddings with mathematical techniques to pinpoint specific timeframes, while the TASK-former model supports hybrid sketch-text search, enabling users to locate scenes using hand-drawn sketches and textual descriptions. Transcript-based search aids users to overcome socio-cultural barriers and OCR-based search utilizes extracted text from video keyframes. The system’s interface allows users to input queries, and browse top-ranked results in a manner that tackles clustered viewing problems seen in common systems. In addition, users can explore visually similar images, and preview short clips, improving both precision and accessibility in video content retrieval.