Contrastive Learning and Temporal Context for Robust Video Identification in Social Networks
摘要
A novel method for video identification is explored using temporal context aggregation and contrastive learning. The proposed approach extracts visual features from video frames using ResNet50, followed by temporal feature aggregation using a Bidirectional Gated Recurrent Unit (GRU). A contrastive learning is introduced to differentiate between similar and dissimilar video pairs, optimizing the model to distinguish unique video identities. The experiments on the UCF101 dataset demonstrates the effectiveness of combining data augmentation, deep feature extraction, and temporal context modeling, yielding valuable insights for video identification in social networks.