Object detection is crucial in many industrial computer vision applications. However, the reliance on manually annotated data prevents a cost-effective deployment of supervised models in domain-specific industrial tasks. This research work presents an evaluation of the real-world performance of leading algorithms in the context of industrial multi-class object detection while being trained solely on synthetic data. We train widely-used models on an order-of-magnitude less synthetic industrial data than the current state-of-the-art and demonstrate a mAP@50-95 of 75% under high-variability environments. Our work also offers an ablation study to narrow the sim-to-real domain gap based on context-aware identification of synthetic features that contribute the most to closing the sim-to-real gap. We show that by employing guided domain randomization based on low-level and on semantic contextual features, it is possible to reduce the amount of required synthetic images by a factor of three while affecting mAP@50-95 on real data by only 2%.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Data-centric Evaluation of Leading Multi-class Object Detection Algorithms Using Synthetic Industrial Data

  • J. Moises Araya-Martinez,
  • Sarvenaz Sardari,
  • Mats Lambert,
  • J. Alexander Zak,
  • Florian Töper,
  • Jörg Krüger,
  • Jens Lambrecht

摘要

Object detection is crucial in many industrial computer vision applications. However, the reliance on manually annotated data prevents a cost-effective deployment of supervised models in domain-specific industrial tasks. This research work presents an evaluation of the real-world performance of leading algorithms in the context of industrial multi-class object detection while being trained solely on synthetic data. We train widely-used models on an order-of-magnitude less synthetic industrial data than the current state-of-the-art and demonstrate a mAP@50-95 of 75% under high-variability environments. Our work also offers an ablation study to narrow the sim-to-real domain gap based on context-aware identification of synthetic features that contribute the most to closing the sim-to-real gap. We show that by employing guided domain randomization based on low-level and on semantic contextual features, it is possible to reduce the amount of required synthetic images by a factor of three while affecting mAP@50-95 on real data by only 2%.