Improving Predictive Accuracy of Multilabel Classification for Cancer Subtypes in Imbalanced Datasets
摘要
The diagnosis of cancer is a difficult task due to symptoms being confused as benign when the time to diagnose is critical in the survival of patients. The situation is compounded by rare cancer types which general practitioners and patients alike can only assume may be possible when all other explanations have been exhausted. Multilabel classification of cancer subtypes in health survey data is challenged by class imbalance and erroneous data. Many prior studies focus on the classification of single cancer types, and health information for avoiding or surviving cancers is typically isolated rather than generalized. Knowledge of chronic illness risk factors can aid healthier lifestyles and assist the recognition of signs requiring prompt medical attention. By transforming survey keywords into linear variables, we merge annual Behavioural and Risk Factor Surveillance System (BRFSS) health surveys, to boost rare cancer subtype samples of data, and enable optimal Synthetic Minority Oversampling Technique (SMOTE) performance. Our iterative method of risk factor and symptom feature selection viz. Alcohol (A), Diabetes (D), Gender + Weight (GW), Gender + Injury + Fatigue (GIF), and Smoking (S) deliver an explainable predictive machine learning (ML) model for many cancer subtypes and cancer occurrences.