<p>Knowledge distillation has recently attracted widespread attention. It is worth noting that masked feature-based distillation is currently a research focus. Currently, research focuses more on improving masked strategies and neglects feature generative strategies. This issue limits the improvement of distillation performance. To address this issue, we design a more effective feature generative framework and propose a multi-scale masked generative distillation method (MMGD). Compared to traditional single scale feature generative methods, we found that using multi-scale feature extraction and fusion methods to capture teacher features of different scales to guide students in generating complete features can enhance the feature expression ability of student models at different scales. In addition, we found that the feature maps that were randomly masked had noise interference between pixels, so a noise interference prevention module (NIPM) was proposed. It consists of special convolution kernels and activation functions. NIPM is applied to the multi-scale feature generative process. It can effectively suppress the noise interference of the masked part of the feature map on the unmasked part during the multi-scale feature extraction and fusion process in MMGD. We conducted a large number of data comparison experiments, ablation experiments, and sensitivity experiments on different models. The experimental results showed that all student models made good progress. Notably, we have increased the accuracy, recall and precision of ResNet-18 in CIFAR100 from 77.09% to 80.20%, from 77.01% to 80.13% and from 77.10% to 80.24%. Compared with previous advanced methods, MMGD also has better performance.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Mmgd:multi-scale masked generative distillation

  • Fang Liu,
  • Haoyu Wang,
  • Weixing Su,
  • Linfeng Li,
  • Tao Huang,
  • Zengxin Guo

摘要

Knowledge distillation has recently attracted widespread attention. It is worth noting that masked feature-based distillation is currently a research focus. Currently, research focuses more on improving masked strategies and neglects feature generative strategies. This issue limits the improvement of distillation performance. To address this issue, we design a more effective feature generative framework and propose a multi-scale masked generative distillation method (MMGD). Compared to traditional single scale feature generative methods, we found that using multi-scale feature extraction and fusion methods to capture teacher features of different scales to guide students in generating complete features can enhance the feature expression ability of student models at different scales. In addition, we found that the feature maps that were randomly masked had noise interference between pixels, so a noise interference prevention module (NIPM) was proposed. It consists of special convolution kernels and activation functions. NIPM is applied to the multi-scale feature generative process. It can effectively suppress the noise interference of the masked part of the feature map on the unmasked part during the multi-scale feature extraction and fusion process in MMGD. We conducted a large number of data comparison experiments, ablation experiments, and sensitivity experiments on different models. The experimental results showed that all student models made good progress. Notably, we have increased the accuracy, recall and precision of ResNet-18 in CIFAR100 from 77.09% to 80.20%, from 77.01% to 80.13% and from 77.10% to 80.24%. Compared with previous advanced methods, MMGD also has better performance.