Robotic systems that utilize machine learning typically demand a large amount of highly accurate data for model training and performance improvement. Crowdsourcing provides researchers with a low-cost opportunity to gather diverse data from around the world. These data can significantly enhance robots’ capabilities in terms of perception, decision-making, and interaction while also accelerating the optimization of machine learning models. However, owing to variations in the skills, experience, and personal biases of crowdsourcing participants, noise and redundancy are often introduced during data collection, compromising data quality and weakening the effectiveness of crowdsourcing data in improving robotic systems. To address this issue, confidence learning (CL) technology was first introduced to filter out erroneous or unreliable crowdsourced data. By analyzing the explicit or implicit relationships between workers and tasks, graph neural networks (GNNs) can be used to leverage node embeddings to infer the true labels of annotated data, improving accuracy and reliability. The experiments were conducted on real crowdsourced annotation datasets, covering 13 real-world application scenarios. The results demonstrate that after applying these processing steps, the accuracy of the data labels is significantly improved, providing a more robust data foundation for training robotic systems.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Crowdsourcing Data Accuracy for Robot Training via Confidence Learning and GNN

  • Wenjun Tang,
  • Rong Chen,
  • Zhikang Zhang,
  • Shikai Guo

摘要

Robotic systems that utilize machine learning typically demand a large amount of highly accurate data for model training and performance improvement. Crowdsourcing provides researchers with a low-cost opportunity to gather diverse data from around the world. These data can significantly enhance robots’ capabilities in terms of perception, decision-making, and interaction while also accelerating the optimization of machine learning models. However, owing to variations in the skills, experience, and personal biases of crowdsourcing participants, noise and redundancy are often introduced during data collection, compromising data quality and weakening the effectiveness of crowdsourcing data in improving robotic systems. To address this issue, confidence learning (CL) technology was first introduced to filter out erroneous or unreliable crowdsourced data. By analyzing the explicit or implicit relationships between workers and tasks, graph neural networks (GNNs) can be used to leverage node embeddings to infer the true labels of annotated data, improving accuracy and reliability. The experiments were conducted on real crowdsourced annotation datasets, covering 13 real-world application scenarios. The results demonstrate that after applying these processing steps, the accuracy of the data labels is significantly improved, providing a more robust data foundation for training robotic systems.