Vision Transformers (ViTs) have garnered significant attention for their superior performance in vision recognition. However, they face two practical challenges: high computational costs and vulnerability to adversarial attacks. To overcome these issues, we propose a novel automatic search framework for adversarially robust and GPU-friendly sparse vision transformers. Our approach uses complexity-aware search to assign different connection patterns for each transformer layer. Additionally, an information bottleneck-driven N:M pruning metric is used to determine which weights to prune in the sparse layers. Experimental results demonstrate that our method reduces parameters by 45.52% to 48.49%, with minimal impact on accuracy and adversarial robustness, making it a practical solution for deploying ViTs in resource-constrained and security-critical scenarios.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

RobSparse: Automatic Search for GPU-Friendly Robust and Sparse Vision Transformers

  • Yulan Su,
  • Sisi Zhang,
  • Yan Wang,
  • Xingbin Wang,
  • Lutan Zhao,
  • Dan Meng,
  • Rui Hou

摘要

Vision Transformers (ViTs) have garnered significant attention for their superior performance in vision recognition. However, they face two practical challenges: high computational costs and vulnerability to adversarial attacks. To overcome these issues, we propose a novel automatic search framework for adversarially robust and GPU-friendly sparse vision transformers. Our approach uses complexity-aware search to assign different connection patterns for each transformer layer. Additionally, an information bottleneck-driven N:M pruning metric is used to determine which weights to prune in the sparse layers. Experimental results demonstrate that our method reduces parameters by 45.52% to 48.49%, with minimal impact on accuracy and adversarial robustness, making it a practical solution for deploying ViTs in resource-constrained and security-critical scenarios.