There are several classification methods in machine learning approaches but the choice of classification method depends on the nature of the data, the complexity of the problem, the interpretability required, and the available resources. Several classification techniques from statistics and machine learning area have been applied to classify cancer data. Therefore, this study aimed to measure the performance of classification methods in supervised machine learning approaches to detect lifestyle factors of cancer patients. A secondary data was used containing 384 patients suffering from different types of cancer in Bangladesh. Seven supervised machine learning methods including Logistic Regression (LR), Linear Discriminate Analysis (LDA), Decision Tree (DT), Support Vector Machine (SVM), Random Forest (RF), Naïve Bayes (NB), and K-Nearest Neighbor classifier (KNN classifier) were used to detect the lifestyle factors of these cancer patients. The target variable was binary in nature, i.e., each cancer type and the input variables were 11 types of food habits of the patients. The performance of the classification methods was measured by using several performance matrices for classification like AUC, sensitivity, specificity, precision, and F1 score. The comparative results of random forest were better than all other techniques considering AUC, Precision, and F1-score for the cancer dataset. For each type of cancer dataset, random forest model provided larger area under the curve (AUC), i.e., more than 0.9 at all times indicating an outstanding performance of the model followed by KNN and SVM. The higher value of precision and F1-score of random forest also indicated that the testing process was working well for random forest method in the dataset. Random forest model also gave the value of mean decrease Gini coefficients for each input variable which indicated which lifestyle factors were contributing more for developing cancers. Random forest is a popular and powerful machine learning approach used in cancer data analysis for both classification and regression tasks. This method has gained popularity due to its high accuracy, ability to handle high-dimensional data, and resistance to over fitting.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Measuring the Performance of Supervised Machine Learning Approaches Using Cancer Data

  • Shahnaj Sultana Sathi,
  • Mohammad Ohid Ullah

摘要

There are several classification methods in machine learning approaches but the choice of classification method depends on the nature of the data, the complexity of the problem, the interpretability required, and the available resources. Several classification techniques from statistics and machine learning area have been applied to classify cancer data. Therefore, this study aimed to measure the performance of classification methods in supervised machine learning approaches to detect lifestyle factors of cancer patients. A secondary data was used containing 384 patients suffering from different types of cancer in Bangladesh. Seven supervised machine learning methods including Logistic Regression (LR), Linear Discriminate Analysis (LDA), Decision Tree (DT), Support Vector Machine (SVM), Random Forest (RF), Naïve Bayes (NB), and K-Nearest Neighbor classifier (KNN classifier) were used to detect the lifestyle factors of these cancer patients. The target variable was binary in nature, i.e., each cancer type and the input variables were 11 types of food habits of the patients. The performance of the classification methods was measured by using several performance matrices for classification like AUC, sensitivity, specificity, precision, and F1 score. The comparative results of random forest were better than all other techniques considering AUC, Precision, and F1-score for the cancer dataset. For each type of cancer dataset, random forest model provided larger area under the curve (AUC), i.e., more than 0.9 at all times indicating an outstanding performance of the model followed by KNN and SVM. The higher value of precision and F1-score of random forest also indicated that the testing process was working well for random forest method in the dataset. Random forest model also gave the value of mean decrease Gini coefficients for each input variable which indicated which lifestyle factors were contributing more for developing cancers. Random forest is a popular and powerful machine learning approach used in cancer data analysis for both classification and regression tasks. This method has gained popularity due to its high accuracy, ability to handle high-dimensional data, and resistance to over fitting.