<p>Vision-language pre-training (VLP) models have significantly accelerated the progress of vision-language models (VLMs) as a pivotal advancement in multimodal artificial intelligence. Despite robust performance under unimodal adversarial attacks, VLP models are still fragile when faced with multimodal adversarial attacks, especially black-box attacks. So focusing on multimodal adversarial transferability is fundamental to driving reliable and secure VLP models. Recent research including Set-level Guidance Attack (SGA) has shown that maximizing cross-modal interactions markedly improves the transferability of multimodal adversarial examples. This method heavily relies on positive samples augmented through data enhancement, which results in strong attack performance on the source model but poor transferability to other target models or tasks. In this paper, we present a method that leverages negative samples to enrich the diversity of adversarial examples, maximizing cross-modal interactions. Furthermore, we developed a multimodal contrastive learning-based method for generating adversarial examples by considering the influence of dual samples (positive and negative samples) during adversarial example generation. The experimental results show that our multimodal adversarial attack method achieves significant transferability in cross-model and cross-task settings. In cross-model transferability, our approach achieves a 100% attack success rate on the source models ALBEF and CLIP-ViT in the image-text retrieval task, while improving the attack success rate for black-box attacks (ALBEF to CLIP-CNN) by 20.71%, compared to SGA.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Contrastive adversarial learning with dual-sample guidance for transferable attacks on vision-language pre-training models

  • Yiming Ren,
  • Yang Xu,
  • Sicong Zhang,
  • Xiaoyao Xie

摘要

Vision-language pre-training (VLP) models have significantly accelerated the progress of vision-language models (VLMs) as a pivotal advancement in multimodal artificial intelligence. Despite robust performance under unimodal adversarial attacks, VLP models are still fragile when faced with multimodal adversarial attacks, especially black-box attacks. So focusing on multimodal adversarial transferability is fundamental to driving reliable and secure VLP models. Recent research including Set-level Guidance Attack (SGA) has shown that maximizing cross-modal interactions markedly improves the transferability of multimodal adversarial examples. This method heavily relies on positive samples augmented through data enhancement, which results in strong attack performance on the source model but poor transferability to other target models or tasks. In this paper, we present a method that leverages negative samples to enrich the diversity of adversarial examples, maximizing cross-modal interactions. Furthermore, we developed a multimodal contrastive learning-based method for generating adversarial examples by considering the influence of dual samples (positive and negative samples) during adversarial example generation. The experimental results show that our multimodal adversarial attack method achieves significant transferability in cross-model and cross-task settings. In cross-model transferability, our approach achieves a 100% attack success rate on the source models ALBEF and CLIP-ViT in the image-text retrieval task, while improving the attack success rate for black-box attacks (ALBEF to CLIP-CNN) by 20.71%, compared to SGA.