\(A^{3}R\) : Vision Language Pre-training by Attentive Alignment and Attentive Reconstruction
摘要
Recently, vision-language models such as ALBEF/CoCa have achieved significant success in various downstream tasks. However, due to the inherent differences between visual and language modalities, they often encounter three challenges: semantics mismatch, fine-grained alignment missing, and trivial masking. To address these issues, we propose to use intrinsic attention scores within the model to guide the cross-modal interaction through attentive alignment and attentive reconstruction. These approaches help the model to focus on the most important information and avoid inefficient processes. Extensive experiments on several downstream tasks demonstrate the effectiveness of the proposed \(A^{3}R\) . Specifically, on the zero-shot Flickr30K retrieval task, \(A^{3}R\) brings a 1.2%/1.8% improvement in top-1 hit accuracy of image-to-text/text-to-image retrieval, using only 76% of the training time and 68% memory compared to the baseline.