<p>Colorectal cancer (CRC) ranks among the most common and deadly cancers globally, responsible for 9.3% of all cancer-related deaths in 2022. Despite advances in treatment, it remains the third most prevalent cancer and the second leading cause of cancer deaths. This underscores the urgent need for cost-effective, noninvasive screening methods, with gene signature-based biomarkers showing potential for early detection. In this study, we applied machine learning techniques to identify candidate biomarker genes for CRC using gene expression data. After preprocessing and normalization, we identified 6,781 differentially expressed genes (DEGs). Using the ReliefF feature selection algorithm, we narrowed these to 43 significant genes linked to important biological processes like guanylate cyclase activity, nucleotide metabolism, and transporter activity, which are critical in CRC development. Further, the CytoHubba tool, applying the MCC algorithm, pinpointed ten hub genes. We trained six machine learning models namely Random Forest (RF), Support vector machine (SVM), Logistic regression (LR), K-Nearest neighbors (KNN), Extreme gradient boosting (XGB), Multi-Layer Perceptron (MLP) on different subsets of the data with the significant DEGs, identifying the top 20 features from each model. By aggregating the unique features from each subset, we identified four candidate biomarker genes AQP8, GUCA2B, OTOP2, and ZG16 that were common between the hub genes and the features selected by the machine learning models. These genes were validated through TCGA data, showing significant downregulation in CRC and association with poor survival, highlighting their potential as diagnostic markers.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Integrating relieff-based feature selection and ensemble machine learning for robust biomarker identification in colorectal cancer

  • Pritam Bera,
  • Subarna Debnath,
  • Chittabrata Mal,
  • Sunil Kanti Mondal

摘要

Colorectal cancer (CRC) ranks among the most common and deadly cancers globally, responsible for 9.3% of all cancer-related deaths in 2022. Despite advances in treatment, it remains the third most prevalent cancer and the second leading cause of cancer deaths. This underscores the urgent need for cost-effective, noninvasive screening methods, with gene signature-based biomarkers showing potential for early detection. In this study, we applied machine learning techniques to identify candidate biomarker genes for CRC using gene expression data. After preprocessing and normalization, we identified 6,781 differentially expressed genes (DEGs). Using the ReliefF feature selection algorithm, we narrowed these to 43 significant genes linked to important biological processes like guanylate cyclase activity, nucleotide metabolism, and transporter activity, which are critical in CRC development. Further, the CytoHubba tool, applying the MCC algorithm, pinpointed ten hub genes. We trained six machine learning models namely Random Forest (RF), Support vector machine (SVM), Logistic regression (LR), K-Nearest neighbors (KNN), Extreme gradient boosting (XGB), Multi-Layer Perceptron (MLP) on different subsets of the data with the significant DEGs, identifying the top 20 features from each model. By aggregating the unique features from each subset, we identified four candidate biomarker genes AQP8, GUCA2B, OTOP2, and ZG16 that were common between the hub genes and the features selected by the machine learning models. These genes were validated through TCGA data, showing significant downregulation in CRC and association with poor survival, highlighting their potential as diagnostic markers.