Enhancing Crowdsourcing Data Accuracy for Robot Training via Confidence Learning and GNN
摘要
Robotic systems that utilize machine learning typically demand a large amount of highly accurate data for model training and performance improvement. Crowdsourcing provides researchers with a low-cost opportunity to gather diverse data from around the world. These data can significantly enhance robots’ capabilities in terms of perception, decision-making, and interaction while also accelerating the optimization of machine learning models. However, owing to variations in the skills, experience, and personal biases of crowdsourcing participants, noise and redundancy are often introduced during data collection, compromising data quality and weakening the effectiveness of crowdsourcing data in improving robotic systems. To address this issue, confidence learning (CL) technology was first introduced to filter out erroneous or unreliable crowdsourced data. By analyzing the explicit or implicit relationships between workers and tasks, graph neural networks (GNNs) can be used to leverage node embeddings to infer the true labels of annotated data, improving accuracy and reliability. The experiments were conducted on real crowdsourced annotation datasets, covering 13 real-world application scenarios. The results demonstrate that after applying these processing steps, the accuracy of the data labels is significantly improved, providing a more robust data foundation for training robotic systems.