In this work, we introduce a low-rank approach based on the truncated Singular Value Decomposition (SVD) technique to make deep neural networks (DNNs) smaller. Our method focuses on reducing the overall physical size of a neural network model without losing much in terms of its accuracy or increasing its error noticeably. Specifically, we applied our technique to a modified MobileNetV2 model, in the context of automatic plant diseases detection. We chose two of the biggest layers of the model and compressed them. Our goal was to see how this compression affects the breakdown of the layers, the error in rebuilding the layers, and how well the modified model performs. The results showed that the smaller model still predicts very accurately, even with the reduced physical size of its layers. By making these layers smaller, our approach offers a practical way to handle the deployment of large neural networks, especially in devices with limited resources, without compromising their effectiveness.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Towards a Low-Rank Approach to Compress Deep Neural Networks

  • M. Liern-García,
  • A. López-García,
  • C. Marco-Detchart,
  • C. Carrascosa

摘要

In this work, we introduce a low-rank approach based on the truncated Singular Value Decomposition (SVD) technique to make deep neural networks (DNNs) smaller. Our method focuses on reducing the overall physical size of a neural network model without losing much in terms of its accuracy or increasing its error noticeably. Specifically, we applied our technique to a modified MobileNetV2 model, in the context of automatic plant diseases detection. We chose two of the biggest layers of the model and compressed them. Our goal was to see how this compression affects the breakdown of the layers, the error in rebuilding the layers, and how well the modified model performs. The results showed that the smaller model still predicts very accurately, even with the reduced physical size of its layers. By making these layers smaller, our approach offers a practical way to handle the deployment of large neural networks, especially in devices with limited resources, without compromising their effectiveness.