Query-Aware Spatiotemporal Transformer-Based Framework for Enhanced Moment Retrieval in Video Surveillance
摘要
In contemporary security applications such as video surveillance, an accurate understanding of video content remains a critical challenge. Existing methods often fail to effectively detect anomalies and retrieve relevant moments from the surveillance footage. This paper introduces the Query-Aware Spatiotemporal Transformer for Moment Retrieval (QuAST-MR), a transformer-based framework that captures the relationship between the query and video content over space and time. QuAST-MR uses self-attention with temporal masking and query-driven features. This allows it to excel in localizing and retrieving pertinent video segments based on user-defined queries. We implemented QuAST-MR using a ResNet-50 backbone for feature extraction, coupled with a self-attention module. This architecture was evaluated on two benchmark datasets, UAV-VID and VIRAT. Extensive evaluations demonstrate the superiority of QuAST-MR over existing approaches. Compared to the state-of-the-art (SOTA), our model achieves a 2.5% improvement in the mean reciprocal rank (MRR) for moment retrieval on UAV-VID.3.12% increase in the normalized discounted cumulative gain (NDCG) for retrieving relevant video segments on VIRAT. By significantly improving the detection accuracy and facilitating efficient moment retrieval based on user queries, QuAST-MR offers substantial value for security personnel. QuAST-MR represents a significant advancement in video comprehension for security applications. Its effectiveness in addressing retrieval challenges makes it a valuable asset in real-world surveillance tasks.