Prostate cancer remains a significant health concern, demanding accurate diagnostic and prognostic tools. In this research study, we address the challenge of imbalanced datasets in prostate cancer classification by employing two prominent data balancing techniques, Synthetic Minority Over-sampling Technique (SMOTE) and Abunaser. We conduct a comprehensive evaluation by applying 20 diverse machine learning algorithms and a novel deep learning approach specifically designed for this task. Our dataset, collected from Kaggle, comprises 10 distinct labels with 100 samples, exhibiting a severe class imbalance. We first investigate the impact of SMOTE and Abunaser on data balance and subsequent classification performance. The results reveal substantial improvements in classification metrics when utilizing Abunaser, with a significant boost in accuracy, precision, recall, and F1-score for both deep learning and machine learning methods. For deep learning, our proposed algorithm achieves noteworthy results when using Abunaser, with an accuracy of 0.9917, precision of 0.9918, recall of 0.9917, and an F1-score of 0.9917. On the other hand, LabelPropagation emerges as the top-performing machine learning algorithm, achieving an accuracy of 0.9741, precision of 0.9743, recall of 0.9741, and an F1-score of 0.9741. These outcomes underscore the effectiveness of Abunaser in mitigating class imbalance and enhancing the classification performance of both deep learning and traditional machine learning models. These findings have the potential to inform the development of more accurate and reliable prostate cancer diagnostic and prognostic tools, ultimately benefiting both patients and healthcare providers.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparative Analysis of Data Balancing Techniques in Prostate Cancer Classification Using Machine Learning and Deep Learning

  • Abedeleilah Salem Ayyad Elmahmuom,
  • Samy S. Abu-Naser

摘要

Prostate cancer remains a significant health concern, demanding accurate diagnostic and prognostic tools. In this research study, we address the challenge of imbalanced datasets in prostate cancer classification by employing two prominent data balancing techniques, Synthetic Minority Over-sampling Technique (SMOTE) and Abunaser. We conduct a comprehensive evaluation by applying 20 diverse machine learning algorithms and a novel deep learning approach specifically designed for this task. Our dataset, collected from Kaggle, comprises 10 distinct labels with 100 samples, exhibiting a severe class imbalance. We first investigate the impact of SMOTE and Abunaser on data balance and subsequent classification performance. The results reveal substantial improvements in classification metrics when utilizing Abunaser, with a significant boost in accuracy, precision, recall, and F1-score for both deep learning and machine learning methods. For deep learning, our proposed algorithm achieves noteworthy results when using Abunaser, with an accuracy of 0.9917, precision of 0.9918, recall of 0.9917, and an F1-score of 0.9917. On the other hand, LabelPropagation emerges as the top-performing machine learning algorithm, achieving an accuracy of 0.9741, precision of 0.9743, recall of 0.9741, and an F1-score of 0.9741. These outcomes underscore the effectiveness of Abunaser in mitigating class imbalance and enhancing the classification performance of both deep learning and traditional machine learning models. These findings have the potential to inform the development of more accurate and reliable prostate cancer diagnostic and prognostic tools, ultimately benefiting both patients and healthcare providers.