<p>In multi-target scenarios, machine vision applications face challenges such as target occlusion, extremely small object sizes, high rates of false alarms and missed detections, as well as ambiguous instructions for manual target selection. To tackle these issues, this paper proposes a Region of Interest-Position Detection (ROI-PD) model designed for human-in-the-loop target localization, which offers distinct advantages in three key aspects: First, in terms of human-in-the-loop interaction, the model incorporates an instruction design that allows users to explicitly select and specify the Region of Interest (ROI), thereby guiding the system to accurately delineate target objects. Second, for enhanced small object detection, we develop an improved Darknet99 feature extraction network by integrating the Squeeze-and-Excitation (SE) attention mechanism with Hybrid Dilated Convolution (HDC). This design significantly boosts the recognition capability for small targets. Experimental results demonstrate that, compared to the original DarkNet53, our DarkNet99 achieves a 2.3% improvement in recognition accuracy while reducing the number of parameters by 28.66% and the computational complexity by 27.37%. Third, highlighting its suitability for edge computing, the proposed model is validated on the Pascal VOC dataset. It exhibits a 3.8% improvement in detection performance (mAP) and a remarkable 34.29% increase in inference speed compared with the YOLOv3-416 model, which clearly underscores its efficiency and practicality for edge deployment.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Human-in-the-loop target location detection method based on the region of interest in multi-object scenes

  • Wenbo Zhang,
  • Zhihui Wang,
  • Xiaoqing Yuan,
  • Zhenhua Yu,
  • Wendong Wang

摘要

In multi-target scenarios, machine vision applications face challenges such as target occlusion, extremely small object sizes, high rates of false alarms and missed detections, as well as ambiguous instructions for manual target selection. To tackle these issues, this paper proposes a Region of Interest-Position Detection (ROI-PD) model designed for human-in-the-loop target localization, which offers distinct advantages in three key aspects: First, in terms of human-in-the-loop interaction, the model incorporates an instruction design that allows users to explicitly select and specify the Region of Interest (ROI), thereby guiding the system to accurately delineate target objects. Second, for enhanced small object detection, we develop an improved Darknet99 feature extraction network by integrating the Squeeze-and-Excitation (SE) attention mechanism with Hybrid Dilated Convolution (HDC). This design significantly boosts the recognition capability for small targets. Experimental results demonstrate that, compared to the original DarkNet53, our DarkNet99 achieves a 2.3% improvement in recognition accuracy while reducing the number of parameters by 28.66% and the computational complexity by 27.37%. Third, highlighting its suitability for edge computing, the proposed model is validated on the Pascal VOC dataset. It exhibits a 3.8% improvement in detection performance (mAP) and a remarkable 34.29% increase in inference speed compared with the YOLOv3-416 model, which clearly underscores its efficiency and practicality for edge deployment.