Hierarchical feature selection based on knowledge and data correlation
摘要
In hierarchical classification learning, data categories exhibit a hierarchical structure. Many studies only have focused on category structure information, knowledge and data correlation are often overlooked. Based on this, a Hierarchical Feature Selection Method Based on Knowledge and Data Correlation (HFSKDC) is proposed. In which, knowledge correlation focuses on maximizing differences between sibling classes, while data correlation considers the feature diversity of categories. First, the sibling strategy is utilized to maximize the differences between categories at the same granularity. Then, data correlation is used to punish samples with different characteristics but consistent labels. Finally, Knowledge and data correlation two regularization terms are unified and optimized into a hierarchical feature selection model. To verify the effectiveness of this proposed method, we conduct experiments on five existing hierarchical classification feature selection methods and eight hierarchical datasets, and the results show that our method is effective and feasible.