<p>Fuzzy rough sets constitute a significant granular computing model in knowledge discovery and have been widely applied to feature selection. However, in heterogeneous and nominal data, most existing fuzzy rough set methods rely on Hamming distance to measure dissimilarity between nominal attribute values, which fails to fully capture their underlying relationships and their impact on decision outcomes. To address this limitation, we propose a fuzzy rough fitting model with nominal distribution metric embedding (FR-NDM). First, the concept of nominal distribution and decision probability is defined, and suitable forms of the nominal distribution metric (NDM) are constructed for diverse data distribution scenarios. Second, a heterogeneous fuzzy information granule with dual-parameter adjustment is developed to accommodate complex data structures. Additionally, fitting approximation operators are established by introducing the judgment condition to ensure that samples attain the maximum membership degree within their respective decision categories. Third, a forward search feature selection algorithm is designed based on FR-NDM. Finally, the proposed method is evaluated on 24 public datasets and compared with 8 state-of-the-art feature selection methods. Experimental results demonstrate the superior performance of our approach.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Feature selection based on fuzzy rough fitting model with nominal distribution metric

  • Jin Qian,
  • Shaowei Yan,
  • Ying Yu,
  • Yongting Ni,
  • Duoqian Miao

摘要

Fuzzy rough sets constitute a significant granular computing model in knowledge discovery and have been widely applied to feature selection. However, in heterogeneous and nominal data, most existing fuzzy rough set methods rely on Hamming distance to measure dissimilarity between nominal attribute values, which fails to fully capture their underlying relationships and their impact on decision outcomes. To address this limitation, we propose a fuzzy rough fitting model with nominal distribution metric embedding (FR-NDM). First, the concept of nominal distribution and decision probability is defined, and suitable forms of the nominal distribution metric (NDM) are constructed for diverse data distribution scenarios. Second, a heterogeneous fuzzy information granule with dual-parameter adjustment is developed to accommodate complex data structures. Additionally, fitting approximation operators are established by introducing the judgment condition to ensure that samples attain the maximum membership degree within their respective decision categories. Third, a forward search feature selection algorithm is designed based on FR-NDM. Finally, the proposed method is evaluated on 24 public datasets and compared with 8 state-of-the-art feature selection methods. Experimental results demonstrate the superior performance of our approach.