A Comprehensive Video Event Retrieval System for Vietnamese News: Integrating CLIP ViT, TASK-Former, Transcripts, and OCR
摘要
In response to the growing need for precise and efficient video retrieval, we present a versatile video event retrieval system that supports multiple query modalities. Our system integrates 5 modes for querying in natural languages: quick search, temporal search, hybrid text and sketch search, transcript-based search, and OCR-based search. Powered by the CLIP ViT-L model, the quick search matches user queries to relevant video segments. Temporal search combines CLIP embeddings with mathematical techniques to pinpoint specific timeframes, while the TASK-former model supports hybrid sketch-text search, enabling users to locate scenes using hand-drawn sketches and textual descriptions. Transcript-based search aids users to overcome socio-cultural barriers and OCR-based search utilizes extracted text from video keyframes. The system’s interface allows users to input queries, and browse top-ranked results in a manner that tackles clustered viewing problems seen in common systems. In addition, users can explore visually similar images, and preview short clips, improving both precision and accessibility in video content retrieval.