<p>Label distribution learning (LDL) is an effective tool to process multi-label data where the label distribution is a probability distribution. Feature selection reduces data dimension, eliminates the impact of irrelevant features and enhances model performance. Rough set theory can be applied for feature selection in a LDL data. However, in most cases this theory only consider lower approximation when it is used to feature selection. In fact, uncertainty of information is related to both the upper and lower approximations. When measuring uncertainty, self information considers both the upper and lower approximations. This paper utilizes self information for feature selection in a LDL data. First of all, distance matrices in the feature space and the label space in a LDL data are constructed, respectively. Then, the upper and lower approximations in a LDL data are proposed. Subsequently, four types of self information (certain decision <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="13042_2025_2771_Article_IEq1.gif" Format="GIF" Height="10" Rendition="HTML" Resolution="72" Type="Linedraw" Width="14" /> </InlineMediaObject> <EquationSource Format="TEX">\(\alpha\)</EquationSource> <EquationSource Format="MATHML"><math> <mi>α</mi> </math></EquationSource> </InlineEquation>-self information, possible decision <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="13042_2025_2771_Article_IEq1.gif" Format="GIF" Height="10" Rendition="HTML" Resolution="72" Type="Linedraw" Width="14" /> </InlineMediaObject> <EquationSource Format="TEX">\(\alpha\)</EquationSource> <EquationSource Format="MATHML"><math> <mi>α</mi> </math></EquationSource> </InlineEquation>-self information, <InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="13042_2025_2771_Article_IEq1.gif" Format="GIF" Height="10" Rendition="HTML" Resolution="72" Type="Linedraw" Width="14" /> </InlineMediaObject> <EquationSource Format="TEX">\(\alpha\)</EquationSource> <EquationSource Format="MATHML"><math> <mi>α</mi> </math></EquationSource> </InlineEquation>-self information and relative <InlineEquation ID="IEq4"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="13042_2025_2771_Article_IEq1.gif" Format="GIF" Height="10" Rendition="HTML" Resolution="72" Type="Linedraw" Width="14" /> </InlineMediaObject> <EquationSource Format="TEX">\(\alpha\)</EquationSource> <EquationSource Format="MATHML"><math> <mi>α</mi> </math></EquationSource> </InlineEquation>-self information) are defined to measure the uncertainty of a LDL data. Next, the best performance of self information: relative <InlineEquation ID="IEq5"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="13042_2025_2771_Article_IEq1.gif" Format="GIF" Height="10" Rendition="HTML" Resolution="72" Type="Linedraw" Width="14" /> </InlineMediaObject> <EquationSource Format="TEX">\(\alpha\)</EquationSource> <EquationSource Format="MATHML"><math> <mi>α</mi> </math></EquationSource> </InlineEquation>-self information is selected by numerical analysis, a feature selection algorithm for a LDL data is designed using the selected self information. Finally, the designed algorithm is tested on 9 standard LDL datasets, and 6 indicators is used in experimental evaluation. The results demonstrate that the designed algorithm has the better performance of classification than 5 excellent feature selection algorithms.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Self information-based feature selection for label distribution learning

  • Zhengwei Zhao,
  • Konglan Huang,
  • Zhaowen Li

摘要

Label distribution learning (LDL) is an effective tool to process multi-label data where the label distribution is a probability distribution. Feature selection reduces data dimension, eliminates the impact of irrelevant features and enhances model performance. Rough set theory can be applied for feature selection in a LDL data. However, in most cases this theory only consider lower approximation when it is used to feature selection. In fact, uncertainty of information is related to both the upper and lower approximations. When measuring uncertainty, self information considers both the upper and lower approximations. This paper utilizes self information for feature selection in a LDL data. First of all, distance matrices in the feature space and the label space in a LDL data are constructed, respectively. Then, the upper and lower approximations in a LDL data are proposed. Subsequently, four types of self information (certain decision \(\alpha\) α -self information, possible decision \(\alpha\) α -self information, \(\alpha\) α -self information and relative \(\alpha\) α -self information) are defined to measure the uncertainty of a LDL data. Next, the best performance of self information: relative \(\alpha\) α -self information is selected by numerical analysis, a feature selection algorithm for a LDL data is designed using the selected self information. Finally, the designed algorithm is tested on 9 standard LDL datasets, and 6 indicators is used in experimental evaluation. The results demonstrate that the designed algorithm has the better performance of classification than 5 excellent feature selection algorithms.