Incorrectly labelled instances can significantly degrade both the performance of the machine learning model and the quality of its assessment. In this work, we describe a system for identifying mislabelled images. On the CIFAR-10 dataset, the achieved \(F_1\) score is higher by ten percentage points compared to another state-of-the-art result. Additionally, a new data augmentation scheme based on the selection of intensities for each transformation and the purity of an image is proposed and applied to the analyzed problem.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Identification of Mislabelled Images with Data Augmentation

  • Jakub Kozieł,
  • Patryk Tomaszewski,
  • Stanisław Kaźmierczak,
  • Jacek Mańdziuk

摘要

Incorrectly labelled instances can significantly degrade both the performance of the machine learning model and the quality of its assessment. In this work, we describe a system for identifying mislabelled images. On the CIFAR-10 dataset, the achieved \(F_1\) score is higher by ten percentage points compared to another state-of-the-art result. Additionally, a new data augmentation scheme based on the selection of intensities for each transformation and the purity of an image is proposed and applied to the analyzed problem.