Feature Selection for Label Distribution Data via Statistical Distribution of Data and Fuzzy Self-Information
摘要
Label distribution learning (LDL) has been extensively applied in diverse fields such as managing label uncertainty, analyzing facial sentiment, and recognizing fuzzy images. In LDL data, each sample is associated with several labels at once, and feature space of samples is with high-dimensionality. As a result, the main challenge in feature selection for a LDL data is to assess the relevance of each feature in relation to a specific set of labels. To address this issue, this paper studies feature selection for a LDL data via statistical distribution of data and fuzzy self information. First, the fuzzy similarity between samples within the feature space is defined, utilizing statistical data distribution and integrating a configurable parameter for similarity adjustment. Through the utilization of different fuzzy similarity radii, the fuzzy similarity relation is established, thereby augmenting the data’s classification efficacy. Then, the decision relation within the label space is presented, and the decision class on each sample is constructed. From this foundation, fuzzy self-information is provided to gauge the uncertainty within a LDL data. Next, a feature selection algorithm for a LDL data is developed based on the selected fuzzy relative self-information. Finally, the experimental results and statistical analysis show the algorithm’s superior on classification performance compared to five advanced feature selection algorithms.