<p>Masked image modeling (MIM) pre-training for large-scale vision transformers (ViTs) has enabled promising downstream performance on top of the learned self-supervised ViT features. In this paper, we question if the <i>extremely simple</i> lightweight ViTs’ fine-tuning performance can also benefit from this pre-training paradigm, which is considerably less studied yet in contrast to the well-established lightweight architecture design methodology. We use an observation-analysis-solution flow for our study. We first systematically <b>observe</b> different behaviors among the evaluated pre-training methods with respect to the downstream fine-tuning data scales. Furthermore, we <b>analyze</b> the layer representation similarities and attention maps across the obtained models, which clearly show the inferior learning of MIM pre-training on higher layers, leading to unsatisfactory transfer performance on data-insufficient downstream tasks. This finding is naturally a guide to designing our distillation strategies during pre-training to <b>solve</b> the above deterioration problem. Extensive experiments have demonstrated the effectiveness of our approach. Our pre-training with distillation on pure lightweight ViTs with vanilla/hierarchical design (5.7<i>M</i>/6.5<i>M</i>) can achieve <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11263_2024_2327_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="43" /> </InlineMediaObject> <EquationSource Format="TEX">\(79.4\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>79.4</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation>/<InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11263_2024_2327_Article_IEq2.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="43" /> </InlineMediaObject> <EquationSource Format="TEX">\(78.9\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>78.9</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation> top-1 accuracy on ImageNet-1K. It also enables SOTA performance on the ADE20K segmentation task (<InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11263_2024_2327_Article_IEq3.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="43" /> </InlineMediaObject> <EquationSource Format="TEX">\(42.8\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>42.8</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation> mIoU) and LaSOT tracking task (<InlineEquation ID="IEq4"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11263_2024_2327_Article_IEq4.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="43" /> </InlineMediaObject> <EquationSource Format="TEX">\(66.1\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>66.1</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation> AUC) in the lightweight regime. The latter even surpasses all the current SOTA lightweight CPU-realtime trackers.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Experimental Study on Exploring Strong Lightweight Vision Transformers via Masked Image Modeling Pre-training

  • Jin Gao,
  • Shubo Lin,
  • Shaoru Wang,
  • Yutong Kou,
  • Zeming Li,
  • Liang Li,
  • Congxuan Zhang,
  • Xiaoqin Zhang,
  • Yizheng Wang,
  • Weiming Hu

摘要

Masked image modeling (MIM) pre-training for large-scale vision transformers (ViTs) has enabled promising downstream performance on top of the learned self-supervised ViT features. In this paper, we question if the extremely simple lightweight ViTs’ fine-tuning performance can also benefit from this pre-training paradigm, which is considerably less studied yet in contrast to the well-established lightweight architecture design methodology. We use an observation-analysis-solution flow for our study. We first systematically observe different behaviors among the evaluated pre-training methods with respect to the downstream fine-tuning data scales. Furthermore, we analyze the layer representation similarities and attention maps across the obtained models, which clearly show the inferior learning of MIM pre-training on higher layers, leading to unsatisfactory transfer performance on data-insufficient downstream tasks. This finding is naturally a guide to designing our distillation strategies during pre-training to solve the above deterioration problem. Extensive experiments have demonstrated the effectiveness of our approach. Our pre-training with distillation on pure lightweight ViTs with vanilla/hierarchical design (5.7M/6.5M) can achieve \(79.4\%\) 79.4 % / \(78.9\%\) 78.9 % top-1 accuracy on ImageNet-1K. It also enables SOTA performance on the ADE20K segmentation task ( \(42.8\%\) 42.8 % mIoU) and LaSOT tracking task ( \(66.1\%\) 66.1 % AUC) in the lightweight regime. The latter even surpasses all the current SOTA lightweight CPU-realtime trackers.