In the research wave of model optimization acceleration, the structured pruning method with filter pruning as the core has been widely used in Convolutional Neural Network (CNN), aiming to reduce the computational complexity and storage requirements of the model. However, the current filter pruning, which form is enough single, ignores the relationship between filters and their mutual influence in the pruning process. Small pruning errors are likely to damage the overall generalization ability of the model, which is difficult to restore the original performance, and the high fine-tuning cost is also unbearable. In this paper, we propose a multi-granularity filter pruning method (MGFP). MGFP optimizes information entropy and weight norm to evaluate the importance of the filter, sets different thresholds and pruning ratios for each layer according to the sensitivity of the filter, and combines different granularity pruning methods such as hard pruning, soft pruning, and strip pruning to minimize the impact on model performance while achieving compression effect. Extensive experiments demonstrate the effectiveness of the MGFP method. On CIFAR-10 and ILSVRC-2012 datasets, the compression effects of model parameters and FLOPs are further enhanced without significant accuracy degradation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-granularity Filter Pruning for Deep Neural Network Compression

  • Quanrun Song,
  • Lin Wang,
  • Changtong Ding,
  • Yu Zhao,
  • Yalong Liu,
  • Shichao Geng

摘要

In the research wave of model optimization acceleration, the structured pruning method with filter pruning as the core has been widely used in Convolutional Neural Network (CNN), aiming to reduce the computational complexity and storage requirements of the model. However, the current filter pruning, which form is enough single, ignores the relationship between filters and their mutual influence in the pruning process. Small pruning errors are likely to damage the overall generalization ability of the model, which is difficult to restore the original performance, and the high fine-tuning cost is also unbearable. In this paper, we propose a multi-granularity filter pruning method (MGFP). MGFP optimizes information entropy and weight norm to evaluate the importance of the filter, sets different thresholds and pruning ratios for each layer according to the sensitivity of the filter, and combines different granularity pruning methods such as hard pruning, soft pruning, and strip pruning to minimize the impact on model performance while achieving compression effect. Extensive experiments demonstrate the effectiveness of the MGFP method. On CIFAR-10 and ILSVRC-2012 datasets, the compression effects of model parameters and FLOPs are further enhanced without significant accuracy degradation.