This paper presents an analysis of Sharpness-Aware Minimization (SAM), a recently introduced efficient optimizer that has demonstrated remarkable improvements in the generalization of deep neural networks. A comprehensive analysis of the asymptotic convergence behaviors of the optimizer when integrated with momentum is proposed. We first show the convergence of the gradient sequence to zero and the topological properties of the set of accumulation points generated by the iterative sequence. Under the assumption of the isolation of stationary points, especially when the function is strongly convex, the convergence of the sequence of iterates is ensured. To validate the practical implications of our analysis, we conduct numerical experiments on classification tasks employing well-known deep learning models, including ResNet18 and ResNet34, with standard datasets CIFAR-10, CIFAR-100, MNIST, and Fashion-MNIST. The numerical results show that, in general, incorporating momentum improves both the training process and testing accuracy for SAM rather than just using standard SGD.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Convergence of Sharpness-Aware Minimization with Momentum

  • Pham Duy Khanh,
  • Hoang-Chau Luong,
  • Boris S. Mordukhovich,
  • Dat Ba Tran,
  • Truc Vo

摘要

This paper presents an analysis of Sharpness-Aware Minimization (SAM), a recently introduced efficient optimizer that has demonstrated remarkable improvements in the generalization of deep neural networks. A comprehensive analysis of the asymptotic convergence behaviors of the optimizer when integrated with momentum is proposed. We first show the convergence of the gradient sequence to zero and the topological properties of the set of accumulation points generated by the iterative sequence. Under the assumption of the isolation of stationary points, especially when the function is strongly convex, the convergence of the sequence of iterates is ensured. To validate the practical implications of our analysis, we conduct numerical experiments on classification tasks employing well-known deep learning models, including ResNet18 and ResNet34, with standard datasets CIFAR-10, CIFAR-100, MNIST, and Fashion-MNIST. The numerical results show that, in general, incorporating momentum improves both the training process and testing accuracy for SAM rather than just using standard SGD.