Vision Transformers (ViTs) with self-attention modules have recently achieved great empirical success in many vision tasks. Due to non-convex interactions across layers, however, the theoretical learning and generalization analysis is non-trivial. Following the framework in Li et al. (A theoretical understanding of shallow vision transformers: Learning, generalization, and sample complexity. In: International Conference on Learning Representations (2023)) and based on a data model characterizing both label-relevant and label-irrelevant tokens, this chapter provides the theoretical analysis of training a shallow ViT, i.e., one self-attention layer followed by a two-layer perceptron, for a classification task. We characterize the sample complexity to achieve a zero generalization error. The sample complexity bound is shown to be positively correlated with the inverse of the fraction of label-relevant tokens, the token noise level, and the initial model error. We also prove that a training process using stochastic gradient descent (SGD) leads to a sparse attention map, which is a formal verification of the general intuition about the success of attention. Moreover, the result indicates that a proper token sparsification can improve the test performance by removing label-irrelevant and/or noisy tokens, including spurious correlations.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Learning and Generalization of Vision Transformers

  • Pin-Yu Chen,
  • Sijia Liu

摘要

Vision Transformers (ViTs) with self-attention modules have recently achieved great empirical success in many vision tasks. Due to non-convex interactions across layers, however, the theoretical learning and generalization analysis is non-trivial. Following the framework in Li et al. (A theoretical understanding of shallow vision transformers: Learning, generalization, and sample complexity. In: International Conference on Learning Representations (2023)) and based on a data model characterizing both label-relevant and label-irrelevant tokens, this chapter provides the theoretical analysis of training a shallow ViT, i.e., one self-attention layer followed by a two-layer perceptron, for a classification task. We characterize the sample complexity to achieve a zero generalization error. The sample complexity bound is shown to be positively correlated with the inverse of the fraction of label-relevant tokens, the token noise level, and the initial model error. We also prove that a training process using stochastic gradient descent (SGD) leads to a sparse attention map, which is a formal verification of the general intuition about the success of attention. Moreover, the result indicates that a proper token sparsification can improve the test performance by removing label-irrelevant and/or noisy tokens, including spurious correlations.