<p>With the expansion of data size and dimension, single filtering and wrapper feature selection methods are limited in dealing with the problem of data redundancy. Therefore, hybrid feature selection algorithms have attracted much attention. In this paper, a hybrid feature selection algorithm combining the Mann–Whitney <i>U</i> test, improved nutcracker optimization, and a football team training algorithm is proposed. First, S0 is generated via the Mann–Whitney test <i>U</i> and the elite solution strategy, and S1 is obtained via dynamic adjustment. Then, the improved nutcracker optimization algorithm is subsequently used to optimize S1 to obtain S2. Finally, the optimal feature subset is determined by the improved soccer team training algorithm acting on S2. After testing on 22 datasets, the classification accuracy of the proposed algorithm is greater than 90% on most datasets, and the dimension reduction rate is less than 0.5%, which has significant advantages in the feature selection of high-dimensional datasets.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A novel hybrid feature selection algorithm based on Mann–Whitney U test and double optimization

  • Xueguang Yang,
  • Yuefeng Zheng,
  • Jing Gan,
  • Yang Lu

摘要

With the expansion of data size and dimension, single filtering and wrapper feature selection methods are limited in dealing with the problem of data redundancy. Therefore, hybrid feature selection algorithms have attracted much attention. In this paper, a hybrid feature selection algorithm combining the Mann–Whitney U test, improved nutcracker optimization, and a football team training algorithm is proposed. First, S0 is generated via the Mann–Whitney test U and the elite solution strategy, and S1 is obtained via dynamic adjustment. Then, the improved nutcracker optimization algorithm is subsequently used to optimize S1 to obtain S2. Finally, the optimal feature subset is determined by the improved soccer team training algorithm acting on S2. After testing on 22 datasets, the classification accuracy of the proposed algorithm is greater than 90% on most datasets, and the dimension reduction rate is less than 0.5%, which has significant advantages in the feature selection of high-dimensional datasets.