Self information-based feature selection for label distribution learning
摘要
Label distribution learning (LDL) is an effective tool to process multi-label data where the label distribution is a probability distribution. Feature selection reduces data dimension, eliminates the impact of irrelevant features and enhances model performance. Rough set theory can be applied for feature selection in a LDL data. However, in most cases this theory only consider lower approximation when it is used to feature selection. In fact, uncertainty of information is related to both the upper and lower approximations. When measuring uncertainty, self information considers both the upper and lower approximations. This paper utilizes self information for feature selection in a LDL data. First of all, distance matrices in the feature space and the label space in a LDL data are constructed, respectively. Then, the upper and lower approximations in a LDL data are proposed. Subsequently, four types of self information (certain decision