Explaining the better generalization of label distribution learning for classification
摘要
Label distribution learning (LDL) has shown advantages over traditional single-label learning (SLL) in many real-world applications, but its superiority has not been theoretically understood. In this paper, we attempt to explain why LDL generalizes better than SLL. Label distribution has rich supervision information such that an LDL method can still choose the sub-optimal label from label distribution even if it neglects the optimal one. In comparison, an SLL method has no information to choose from when it fails to predict the optimal label. The better generalization of LDL can be credited to the rich information of label distribution. We further establish the label distribution margin theory to prove this explanation; inspired by the theory, we put forward a novel LDL approach called LDL-LDML. In the experiments, the LDL baselines outperform the SLL ones, and LDL-LDML achieves competitive performance against existing LDL methods, which support our explanation and theories in this paper.