<p>With the rapid advancement of high-performance computing and deep neural networks, video recognition models have achieved significant success, while they remain vulnerable to carefully crafted perturbations. Recent studies have begun to explore the transferability of adversarial attacks across video recognition models. However, current attack methods still rely on the sign function for gradient estimation, which introduces non-differentiability and coarse updates, potentially leading to suboptimal and non-transferable adversarial perturbations. To address this issue, we propose to use more refined perturbations to enhance the transferability of video adversarial examples. Firstly, we analyze the limitations of the commonly used sign function in existing methods. Subsequently, we propose a scaling function to replace the sign function, which computes an adaptive scaling factor based on the <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(L_{1}\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mi>L</mi> <mn>1</mn> </msub> </math></EquationSource> </InlineEquation> norm of the gradient at each iteration, thereby preserving a more effective optimization direction. Finally, we propose a refined perturbations attack (RPA) algorithm tailored for video adversarial attacks. Notably, the proposed scaling function can also be integrated into other attack methods to significantly enhance their transferability. Extensive experiments conducted on the supercomputing cluster demonstrate RPA’s superiority over state-of-the-art methods. For transfer-based black-box attacks against video recognition models, RPA achieves an average attack success rate of 69.40% on Kinetics-400 and 52.64% on UCF-101.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Boosting video adversarial attacks with refined perturbations

  • Xueqiang Han,
  • Shi Wang,
  • Qing Tian,
  • Shihui Zhang

摘要

With the rapid advancement of high-performance computing and deep neural networks, video recognition models have achieved significant success, while they remain vulnerable to carefully crafted perturbations. Recent studies have begun to explore the transferability of adversarial attacks across video recognition models. However, current attack methods still rely on the sign function for gradient estimation, which introduces non-differentiability and coarse updates, potentially leading to suboptimal and non-transferable adversarial perturbations. To address this issue, we propose to use more refined perturbations to enhance the transferability of video adversarial examples. Firstly, we analyze the limitations of the commonly used sign function in existing methods. Subsequently, we propose a scaling function to replace the sign function, which computes an adaptive scaling factor based on the \(L_{1}\) L 1 norm of the gradient at each iteration, thereby preserving a more effective optimization direction. Finally, we propose a refined perturbations attack (RPA) algorithm tailored for video adversarial attacks. Notably, the proposed scaling function can also be integrated into other attack methods to significantly enhance their transferability. Extensive experiments conducted on the supercomputing cluster demonstrate RPA’s superiority over state-of-the-art methods. For transfer-based black-box attacks against video recognition models, RPA achieves an average attack success rate of 69.40% on Kinetics-400 and 52.64% on UCF-101.