Data-driven identification of modifiable risk factors for colorectal cancer in Palestine using machine learning methods
摘要
Colorectal cancer (CRC) is a growing public health concern in Palestine, yet limited evidence is available on locally relevant behavioral, dietary, socioeconomic, and contextual predictors. Machine learning (ML) may support prevention-oriented risk stratification by identifying complex predictive patterns among potentially modifiable factors. This study aimed to identify important predictors of CRC among Palestinian adults and evaluate ML models for CRC classification after excluding screening-related variables.
MethodsA case-control dataset of 216 participants was analyzed, including 108 CRC cases and 108 cancer-free controls. Initially, 57 variables were reviewed, and 39 candidate predictors were retained after excluding non-modifiable variables and the outcome. Screening-related variables, including previous colonoscopy and fecal occult blood testing, were excluded from the primary modeling analysis because they may reflect diagnostic workup, screening access, or healthcare utilization. Multiple feature-selection methods were evaluated, including filter, wrapper, embedded, and hybrid approaches. Ten optimized ML classifiers were trained and internally validated using stratified 5-fold nested cross-validation, with preprocessing, feature selection, hyperparameter optimization, and model training performed within the training folds. Model performance was assessed using accuracy, AUC, sensitivity, specificity, F1-score, and Matthews correlation coefficient. SHAP analysis was used to interpret the optimized Ensemble models following hybrid feature selection.
ResultsAfter excluding screening-related variables, the best model-combination performance was achieved by Forward Selection with the Ensemble classifier, with an accuracy of 90.3% ± 4.3, a balanced accuracy of 90.28%, a sensitivity of 89.81% ± 5.59, and a specificity of 90.74% ± 4.18. Hybrid feature-selection approaches also showed competitive performance with Ensemble learning. Hybrid LASSO + GA with Ensemble and Hybrid mRMR + PSO with Ensemble both achieved a balanced accuracy of 89.81%. SHAP analysis consistently identified physical activity as the most influential predictor, followed by dietary, socioeconomic, and contextual variables, including fruit consumption frequency, yearly income, occupation, and place of residence.
ConclusionsAfter excluding screening-related variables, ML models that relied on lifestyle, socioeconomic, and contextual predictors showed promising internal validation performance in classifying CRC among Palestinian adults. Hybrid feature-selection methods combined with Ensemble learning represent promising candidate pipelines for future validation. However, the findings remain exploratory and require external validation in larger, independent cohorts before clinical or public health implementation.