This study investigates the challenges of creating reliable bounding-box annotations for computer vision datasets through crowdsourcing and proposes the Single Human Rating (SHR) method to address these challenges. The SHR method utilizes an autoregressive (AR) model to predict annotator reliability over time, relying on a single annotator per sample and avoiding iterative corrections or consensus-building approaches. In this study, a baseline dataset was compiled by a machine learning researcher and compared to annotations provided by three collaborator groups with limited machine learning expertise. The results revealed significant disparities in annotation quality, particularly in bounding-box accuracy, between the expert and crowdsourced annotations. SHR was found to be highly effective in predicting reliability across both highly reliable (0.973 and 0.970) and less reliable (0.679) contributors, achieving low root-mean-square error (RMSE) values of 0.022, 0.004, and 0.237 respectively. Additionally, SHR demonstrated robustness in the early detection of declining annotation quality, enabling timely intervention. The results emphasize the importance of adapting quality control measures to annotator reliability levels, providing valuable insights for improving the overall quality of crowdsourced datasets in complex computer vision applications.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Reliability Challenges in Crowdsourced Bounding-Box Data for Computer Vision

  • Fatma Gumus

摘要

This study investigates the challenges of creating reliable bounding-box annotations for computer vision datasets through crowdsourcing and proposes the Single Human Rating (SHR) method to address these challenges. The SHR method utilizes an autoregressive (AR) model to predict annotator reliability over time, relying on a single annotator per sample and avoiding iterative corrections or consensus-building approaches. In this study, a baseline dataset was compiled by a machine learning researcher and compared to annotations provided by three collaborator groups with limited machine learning expertise. The results revealed significant disparities in annotation quality, particularly in bounding-box accuracy, between the expert and crowdsourced annotations. SHR was found to be highly effective in predicting reliability across both highly reliable (0.973 and 0.970) and less reliable (0.679) contributors, achieving low root-mean-square error (RMSE) values of 0.022, 0.004, and 0.237 respectively. Additionally, SHR demonstrated robustness in the early detection of declining annotation quality, enabling timely intervention. The results emphasize the importance of adapting quality control measures to annotator reliability levels, providing valuable insights for improving the overall quality of crowdsourced datasets in complex computer vision applications.