This article proposes a machine learning framework for risk assessment of breast cancer using factors such as those related to the patient’s history, anthropometric measures, clinical symptoms, etc. Breast cancer risk factor dataset including 14,382 cases from Breast Cancer Surveillance Consortium dataset was used in the experiments. The dataset consisted of a healthy group and a control group consisting of patients with breast cancer detected within a year of the screening mammography. Statistical analysis using Mann–Whitney U-test and classification using various classifiers such as support vector machine (SVM) different kernels, Gaussian (fine, medium, and course), Naïve Bayes, k-nearest neighbor (k = 5), and linear discriminant analysis (LDA) were applied on this data set for risk estimation of breast cancer under fivefold data division protocol. Experimental results show that the SVM classifier with cubic kernel achieves the superior classification accuracy of 99.43% with area under curve (AUC) equal to 1.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Machine Learning-Based Intelligent Framework for Breast Cancer Risk Assessment

  • Bikesh Kumar Singh,
  • Narendra Kuber Bodhey,
  • Yogesh Sharma

摘要

This article proposes a machine learning framework for risk assessment of breast cancer using factors such as those related to the patient’s history, anthropometric measures, clinical symptoms, etc. Breast cancer risk factor dataset including 14,382 cases from Breast Cancer Surveillance Consortium dataset was used in the experiments. The dataset consisted of a healthy group and a control group consisting of patients with breast cancer detected within a year of the screening mammography. Statistical analysis using Mann–Whitney U-test and classification using various classifiers such as support vector machine (SVM) different kernels, Gaussian (fine, medium, and course), Naïve Bayes, k-nearest neighbor (k = 5), and linear discriminant analysis (LDA) were applied on this data set for risk estimation of breast cancer under fivefold data division protocol. Experimental results show that the SVM classifier with cubic kernel achieves the superior classification accuracy of 99.43% with area under curve (AUC) equal to 1.