MALIAN: Multi-grained Alignment via Learnable Interaction and Aggregation Normalization for Text-Video Retrieval
摘要
Multi-grained alignment enables various ways to explore the matching relationship more comprehensively and deeply between text and video, which is essential for the text-video retrieval (TVR) task. However, existing multi-grained alignment methods ignore the in-depth utilization of fine-grained visual representation (patch) and the impact of video noise, leading to insufficiently accurate similarity calculations. To this end, we propose a novel Multi-grained Alignment method via Learnable Interaction and Aggregation Normalization, dubbed MALIAN. Specifically, we design a learnable interaction module that first uses multiple learnable tokens to comprehensively learn fine-grained visual representation, then constructs cross-granularity alignment to reduce the influence of redundant information in learnable tokens and finally captures fine-grained and coarse-grained alignment information. Furthermore, we design an aggregation normalization module that considers the importance of different information in both intra- and inter-modalities when aggregating different similarity matrices into a similarity score for each granularity and introduces the optimal transport theory to normalize the similarity scores, thereby alleviating retrieval errors caused by video noise. The outcomes of our experiments reveal that the proposed MALIAN attains new state-of-the-art performance on the MSR-VTT (51.6%), MSVD (48.9%), and Activity-Net (47.6%) datasets.