From Omics Data to Candidate Genes: An Innovative Machine Learning Approach for Biomarker Identification
摘要
Advances in high-throughput technologies have accelerated omics data research. The recent explosion of omics, namely transcriptomic data, has opened new opportunities for the discovery of novel biomarkers with potential to be incorporated into clinical practice. However, due to their extreme complexity, gaining useful insights is particularly challenging. Hence, the application of machine learning techniques on transcriptomic data emerges as a highly promising area for the discovery of new biomarkers. For exploring the potential of these techniques, this paper proposes a novel approach to process gene expression data with the aim of finding candidate gene signatures. Our methodology consists of an ensemble feature selection strategy based on the Boruta, SVM-RFE and LASSO methods, complemented by a second feature selection based on the gene importance calculated by the Random Forest, XGBoost, Support Vector Machine, Logistic Regression and AdaBoost methods. Performing simulations with a dataset of atopic dermatitis patients, our proposal resulted in an 8-gene signature with high AUC (0.839) and accuracy (0.8462) values.