The manufacturing defects of flat aluminum sheet products can affect their visual quality. These defects are difficult to detect due to their small size and irregular shape. In this paper, a novel approach named MSN-YOLOv5 is proposed for detecting small object defects based on YOLOv5. The proposed method combines Multi-Dconv Head Transposed Attention (MDTA), Shuffle Attention (SA), and Normalized Wasserstein Distance (NWD). To enhance the feature extraction network, the advantages of Bottleneck Transformers and MDTA are leveraged, particularly in improving the last C3 module, which is named MD3. This enhancement strengthens the capture of global information. Additionally, the SA module is introduced before the 40x40 detection head of the prediction network to reduce interference from irrelevant background information. Furthermore, the NWD loss function is incorporated, which combines CIOU and NWD weighting, to address the sensitivity of previous Intersection over Union (IoU) loss functions to small object position deviations. The experimental results indicate that the model size is 25.79MB, achieving a of 83.1% and an FPS of 75.1. Compared to the baseline model, the proposed method demonstrates an increase of 4.133% in and 11.4% in FPS.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Small Target Defects Detection of Aluminum Plates Surface Using an MSN-YOLOv5 Model

  • Jianguo Zhang,
  • Jiangwei You,
  • Jianfang Jia,
  • Wenwen Zhang,
  • Xiaoqing Ren

摘要

The manufacturing defects of flat aluminum sheet products can affect their visual quality. These defects are difficult to detect due to their small size and irregular shape. In this paper, a novel approach named MSN-YOLOv5 is proposed for detecting small object defects based on YOLOv5. The proposed method combines Multi-Dconv Head Transposed Attention (MDTA), Shuffle Attention (SA), and Normalized Wasserstein Distance (NWD). To enhance the feature extraction network, the advantages of Bottleneck Transformers and MDTA are leveraged, particularly in improving the last C3 module, which is named MD3. This enhancement strengthens the capture of global information. Additionally, the SA module is introduced before the 40x40 detection head of the prediction network to reduce interference from irrelevant background information. Furthermore, the NWD loss function is incorporated, which combines CIOU and NWD weighting, to address the sensitivity of previous Intersection over Union (IoU) loss functions to small object position deviations. The experimental results indicate that the model size is 25.79MB, achieving a of 83.1% and an FPS of 75.1. Compared to the baseline model, the proposed method demonstrates an increase of 4.133% in and 11.4% in FPS.