<p>Deep learning-based real-time instance segmentation seeks to achieve target detection, recognition, and pixel-level segmentation in video streams or continuous images with minimal latency. However, due to factors such as target scale variation and background noise interference, the precision of current mainstream segmentation models remains insufficient. To mitigate this issue, this paper proposes a novel real-time instance segmentation model named GCAM-Inst. Specifically, we first introduce a Multi-scale Spatial Pyramid Pooling (MSPP) module into the YOLOv10-seg framework to augment the feature representation Competence of the backbone network. Secondly, we propose a Global Coordinate Attention Mechanism (GCAM) to optimize the information processing approach of the model, thereby improving the segmentation performance for small-scale targets. Finally, we optimize the original localization loss in the baseline model and introduce a novel IoU loss metric (OIoU) to enhance the model’s perception of instance locations. GCAM-Inst achieves competitive outcomes derived from two extensive publicly available datasets, MS COCO 2017 and KINS. Compared with the baseline model YOLOv10-seg, GCAM-Inst enhances the average precision (AP) by 2.4% and 2.8% for the two respective datasets.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

GCAM-Inst: real-time instance segmentation via global contextual modeling

  • Chengang Dong,
  • Yongkang Ding,
  • Jianwei Hu

摘要

Deep learning-based real-time instance segmentation seeks to achieve target detection, recognition, and pixel-level segmentation in video streams or continuous images with minimal latency. However, due to factors such as target scale variation and background noise interference, the precision of current mainstream segmentation models remains insufficient. To mitigate this issue, this paper proposes a novel real-time instance segmentation model named GCAM-Inst. Specifically, we first introduce a Multi-scale Spatial Pyramid Pooling (MSPP) module into the YOLOv10-seg framework to augment the feature representation Competence of the backbone network. Secondly, we propose a Global Coordinate Attention Mechanism (GCAM) to optimize the information processing approach of the model, thereby improving the segmentation performance for small-scale targets. Finally, we optimize the original localization loss in the baseline model and introduce a novel IoU loss metric (OIoU) to enhance the model’s perception of instance locations. GCAM-Inst achieves competitive outcomes derived from two extensive publicly available datasets, MS COCO 2017 and KINS. Compared with the baseline model YOLOv10-seg, GCAM-Inst enhances the average precision (AP) by 2.4% and 2.8% for the two respective datasets.