In the domain of image annotation, the involvement of human annotators presents a series of intricate challenges tied to the complexities of visual perception. The manual labeling process demands an understanding of context and visual intricacies, all susceptible to human subjectivity. Furthermore, the presence of visual distractions compromises annotation quality, not only impeding annotation precision but also escalating time expenditures as annotators navigate through visual noise to discern pertinent details. Traditional image annotation pipelines underscore these challenges in favor of automatic or semi-automatic annotation, emphasizing the critical necessity for innovative approaches in annotation tasks where the human annotator role is fundamental. Within this context, the Grounded SAM model, emerges as a potent tool for text-prompt-based panoptic segmentation. This paper proposes a novel annotation pipeline employing Grounded SAM and LaMa cleaner models to augment the indispensable role of human annotators by enhancing annotation efficiency through natural language-based attention mining for visual distractions elimination and preannotation techniques. The effectiveness of distraction elimination is demonstrated through an annotation task involving human annotators, with half of the images processed through our pipeline and the remaining unmodified. With our current approach and with the current data analyzed, image annotation times of 70% of the annotators were reduced by 15.88%, while global annotation time was reduced by a 6.93%.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Image Annotation Through Attention Mining: A Grounded SAM Approach

  • Jaime Cortón-González,
  • Ángel Mora-Sánchez,
  • Silvia Rodríguez-Jiménez

摘要

In the domain of image annotation, the involvement of human annotators presents a series of intricate challenges tied to the complexities of visual perception. The manual labeling process demands an understanding of context and visual intricacies, all susceptible to human subjectivity. Furthermore, the presence of visual distractions compromises annotation quality, not only impeding annotation precision but also escalating time expenditures as annotators navigate through visual noise to discern pertinent details. Traditional image annotation pipelines underscore these challenges in favor of automatic or semi-automatic annotation, emphasizing the critical necessity for innovative approaches in annotation tasks where the human annotator role is fundamental. Within this context, the Grounded SAM model, emerges as a potent tool for text-prompt-based panoptic segmentation. This paper proposes a novel annotation pipeline employing Grounded SAM and LaMa cleaner models to augment the indispensable role of human annotators by enhancing annotation efficiency through natural language-based attention mining for visual distractions elimination and preannotation techniques. The effectiveness of distraction elimination is demonstrated through an annotation task involving human annotators, with half of the images processed through our pipeline and the remaining unmodified. With our current approach and with the current data analyzed, image annotation times of 70% of the annotators were reduced by 15.88%, while global annotation time was reduced by a 6.93%.