<p>Image classification is one of the typical applications of convolutional neural networks, and recently, many models have been proposed to improve classification performance. Particularly, the attention mechanism that simulated the human visual system has attracted many researchers, and several attention mechanisms have been studied to improve the accuracy of image classification and object detection. Typical attention mechanisms used in image classification are spatial attention, channel attention, and spatial and channel attention. These mechanisms apply global mean pooling respectively in spatial and channel dimensions to compress spatial and channel dimensions, which leads information loss and does not take into account the correlation of channel dimensions and spatial dimensions. The importance of spatial information varies in channel, and spatial information and channel information are interdependent. Therefore, we propose a Depthwise Convolution-based Spatial Channel Attention(DwConv-SCA) mechanism to capture spatial attention information across the channel and the interaction information between channel dimension and spatial dimension, which simultaneously captures important spatial information for each channel of the feature map, while taking the interaction between spatial dimension and channel dimension. Also, the proposed attention mechanism can be integrated with existing ResNet, MobileNet, and EfficientNet, and the image classification performance is improved by up to 1% compared to the existing methods, with the Top-1 classification accuracy of 73.4% and 87.9%, respectively, on the ImageNet-1&#xa0;K and PlantVillage datasets.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Design of spatial channel attention mechanism based on depthwise convolution

  • Suk-Hyang Ri,
  • Kwang-Chol Ri

摘要

Image classification is one of the typical applications of convolutional neural networks, and recently, many models have been proposed to improve classification performance. Particularly, the attention mechanism that simulated the human visual system has attracted many researchers, and several attention mechanisms have been studied to improve the accuracy of image classification and object detection. Typical attention mechanisms used in image classification are spatial attention, channel attention, and spatial and channel attention. These mechanisms apply global mean pooling respectively in spatial and channel dimensions to compress spatial and channel dimensions, which leads information loss and does not take into account the correlation of channel dimensions and spatial dimensions. The importance of spatial information varies in channel, and spatial information and channel information are interdependent. Therefore, we propose a Depthwise Convolution-based Spatial Channel Attention(DwConv-SCA) mechanism to capture spatial attention information across the channel and the interaction information between channel dimension and spatial dimension, which simultaneously captures important spatial information for each channel of the feature map, while taking the interaction between spatial dimension and channel dimension. Also, the proposed attention mechanism can be integrated with existing ResNet, MobileNet, and EfficientNet, and the image classification performance is improved by up to 1% compared to the existing methods, with the Top-1 classification accuracy of 73.4% and 87.9%, respectively, on the ImageNet-1 K and PlantVillage datasets.