Boosting video adversarial attacks with refined perturbations
摘要
With the rapid advancement of high-performance computing and deep neural networks, video recognition models have achieved significant success, while they remain vulnerable to carefully crafted perturbations. Recent studies have begun to explore the transferability of adversarial attacks across video recognition models. However, current attack methods still rely on the sign function for gradient estimation, which introduces non-differentiability and coarse updates, potentially leading to suboptimal and non-transferable adversarial perturbations. To address this issue, we propose to use more refined perturbations to enhance the transferability of video adversarial examples. Firstly, we analyze the limitations of the commonly used sign function in existing methods. Subsequently, we propose a scaling function to replace the sign function, which computes an adaptive scaling factor based on the