<p>Onboard cameras of coal mining machinery (e.g., roadheaders) in underground mines often suffer from severe video jitter under complex vibrational conditions, leading to inaccurate tray target recognition. To address this issue, this paper proposes a novel detection method integrating video stabilization and a Squeeze-and-Excitation-Swish (SES) attention mechanism. The proposed approach begins with a video stabilization module that effectively mitigates image blurring and distortion, reducing horizontal and vertical frame offsets by 88.4% and 80.6%, respectively, thereby providing clearer and more stable image sequences for subsequent processing. Subsequently, an SES attention module is embedded into the YOLOv5 architecture to enhance the model’s focus on small target regions via dynamic channel-wise weight recalibration. Experimental results demonstrate that the SES-YOLOv5 algorithm achieves a mAP of 97.1%, outperforming the baseline YOLOv5s by 3.0%. It also maintains an inference speed of 93 FPS, which is 4.5% faster than CBAM-YOLOv5. Furthermore, on various jittery video datasets, the proposed method improves tray recognition accuracy by up to 14% compared to YOLOv8. The proposed method significantly enhances the accuracy and robustness of tray detection in vibrational environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Vibration environment target detection method based on video stabilization and SES attention mechanism

  • Chengcheng Li

摘要

Onboard cameras of coal mining machinery (e.g., roadheaders) in underground mines often suffer from severe video jitter under complex vibrational conditions, leading to inaccurate tray target recognition. To address this issue, this paper proposes a novel detection method integrating video stabilization and a Squeeze-and-Excitation-Swish (SES) attention mechanism. The proposed approach begins with a video stabilization module that effectively mitigates image blurring and distortion, reducing horizontal and vertical frame offsets by 88.4% and 80.6%, respectively, thereby providing clearer and more stable image sequences for subsequent processing. Subsequently, an SES attention module is embedded into the YOLOv5 architecture to enhance the model’s focus on small target regions via dynamic channel-wise weight recalibration. Experimental results demonstrate that the SES-YOLOv5 algorithm achieves a mAP of 97.1%, outperforming the baseline YOLOv5s by 3.0%. It also maintains an inference speed of 93 FPS, which is 4.5% faster than CBAM-YOLOv5. Furthermore, on various jittery video datasets, the proposed method improves tray recognition accuracy by up to 14% compared to YOLOv8. The proposed method significantly enhances the accuracy and robustness of tray detection in vibrational environments.