An interpretable sample selection framework against numerical label noise
摘要
Numerical label noise in regression would misguide the model training and worsen the generalization performance. As a popular technique, noise filtering reduces the noise level by removing mislabeled samples. Some filters care about the noise level so much that a few clean samples are also removed. The existing optimal sample selection (OSS) framework balances the number of removals and the noise level to avoid overcleaning. However, its underlying interpretability is unobvious due to the complicated objective function, and inaccurate parameter estimates or settings may discount the filtering effect. To address these issues, we first propose a novel interpretable sample selection (ISS) framework against numerical label noise. It seeks to maximize the number of available samples while having a relatively low noise level. ISS converges to OSS in