Thyroid cancer risk classification is critical for timely and effective treatment planning. This study explores the application of machine learning (ML) and deep learning (DL) techniques to classify thyroid cancer risk into three categories: low, intermediate, and high. The dataset, obtained from Kaggle, consists of 383 samples with 17 clinical features, including age, gender, smoking history, thyroid function, and pathology. To address class imbalance (248 low-risk, 102 intermediate-risk, 33 high-risk), the Abunaser technique was employed to augment the dataset, resulting in a balanced set of 1500 samples (500 per class). Eleven machine learning models were implemented, including Random Forest Classifier, XGBoost Classifier, Logistic Regression), among others. Additionally, a custom deep learning model was trained for 60 epochs. The dataset was split into training (80%), validation (10%), and testing (10%) sets. The models were evaluated using accuracy, F1-score, recall, and precision. The Random Forest Classifier achieved the best performance among the ML models, with an accuracy of 98.43%, recall of 98.38%, precision of 98.36%, and an F1-score of 98.32%. The deep learning model outperformed the ML models, achieving an accuracy of 99.20%, recall of 99.17%, precision of 99.15%, and an F1-score of 99.14%. These results demonstrate that the proposed deep learning model provides superior performance in predicting thyroid cancer risk, particularly in scenarios with balanced data. This study underscores the potential of deep learning models in improving classification accuracy and highlights the importance of data balancing techniques for enhancing model performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Thyroid Cancer Risk Classification Using Machine Learning and Deep Learning Techniques: A Comparative Study with Balanced Dataset Augmentation

  • Mohammed A. Alkahlout,
  • Samy S. Abu-Naser

摘要

Thyroid cancer risk classification is critical for timely and effective treatment planning. This study explores the application of machine learning (ML) and deep learning (DL) techniques to classify thyroid cancer risk into three categories: low, intermediate, and high. The dataset, obtained from Kaggle, consists of 383 samples with 17 clinical features, including age, gender, smoking history, thyroid function, and pathology. To address class imbalance (248 low-risk, 102 intermediate-risk, 33 high-risk), the Abunaser technique was employed to augment the dataset, resulting in a balanced set of 1500 samples (500 per class). Eleven machine learning models were implemented, including Random Forest Classifier, XGBoost Classifier, Logistic Regression), among others. Additionally, a custom deep learning model was trained for 60 epochs. The dataset was split into training (80%), validation (10%), and testing (10%) sets. The models were evaluated using accuracy, F1-score, recall, and precision. The Random Forest Classifier achieved the best performance among the ML models, with an accuracy of 98.43%, recall of 98.38%, precision of 98.36%, and an F1-score of 98.32%. The deep learning model outperformed the ML models, achieving an accuracy of 99.20%, recall of 99.17%, precision of 99.15%, and an F1-score of 99.14%. These results demonstrate that the proposed deep learning model provides superior performance in predicting thyroid cancer risk, particularly in scenarios with balanced data. This study underscores the potential of deep learning models in improving classification accuracy and highlights the importance of data balancing techniques for enhancing model performance.