Ensemble filter RL feature selection method based on MW-shaped transfer function for high-dimensional cancer gene expression data
摘要
Cancer gene expression data present challenges such as high dimensionality, multiple samples, and multi-class classification, making the task of cancer subtype diagnosis complex and computationally demanding. Feature selection (FS) plays a crucial role in addressing these challenges by reducing the dimensionality of the data while retaining the most relevant features, thus improving classification performance. In this paper, a novel MW-shaped transfer function was proposed, which is derived from the combination of M-shaped and W-shaped transfer functions. This transfer function is integrated with two parallel filtering algorithms, ReliefF and Local Learning-based Clustering (LLC), to form an Ensemble Filter RL method. The goal of the proposed approach is to enhance classification accuracy by adaptively selecting and removing features based on their relevance to the target classification task. To validate the effectiveness of the MW-shaped transfer function and the Ensemble Filter RL method, experiments were conducted on 12 high-dimensional UCI datasets, where the Pelican Optimization Algorithm (POA) was utilized to select the optimal transfer function. In addition, further validation was performed on 12 cancer gene expression datasets, which were selected to assess the performance of the proposed method in dealing with high-dimensional gene expression data. These datasets include a variety of cancer types and provide a robust evaluation framework for the proposed method’s ability to handle complex, high-dimensional biological data. The experimental results demonstrate that the proposed approach significantly reduces the number of selected features, improves classification accuracy, and achieves lower fitness values when compared to other optimization algorithms. These results highlight the robustness and effectiveness of the MW-shaped transfer function and Ensemble Filter RL in cancer gene expression data classification, offering a promising solution for high-dimensional feature selection in bioinformatics applications.