Colorectal cancer remains a significant health challenge, necessitating the implementation of efficient methods for early identification. The application of artificial intelligence, particularly deep learning, for tissue classification holds the potential to improve diagnostic accuracy and drive breakthroughs in clinical oncology techniques. This study presents a machine learning approach using the features extracted from a pre-trained deep learning model to classify nine classes of human colorectal cancer (CRC) and normal tissue. The proposed method is trained and evaluated on the public colorectal cancer histopathological images (NCT-CRC-HE-100K and CRC-VAL-HE-7K). The training set (NCT-CRC-HE-100K) contains 100,000 histopathological images and the testing set (CRC-VAL-HE-7K) has 7180 images of nine classes. In this study, the smaller training set is created by sampling only 10,000 images of nine classes (10% of the original training set) and keeping the same ratio of these classes as the original training set. The pre-trained ResNet18 is applied to extract a vector of 512 features for each image and then the linear discriminant analysis (LDA) transformation reduces the data dimensions for the new training set and the original testing set. After that, five machine learning, i.e., Random forest (RF), Support vector machine (SVM), Light gradient boosting machine (LightGBM), Categorical gradient boosting (CatBoost) and Ensemble model (utilizing weighted average technique with the four above methods) are trained and evaluated on these sets. Furthermore, random search cross-validation technique is applied to all classifiers for finding the optimal hyperparameters. The overall accuracy of Ensemble model is highest at 89.8% and SVM, RF, LightGBM and CatBoost obtain 89.6%, 89.4%, 89.7% and 89.4% on testing set 1, respectively. Meanwhile, SVM classifier achieves the highest accuracy of 93.3% and four other classifiers (i.e., RF, LightGBM, CatBoost and Ensemble) have the accuracy of 93.0%, 93.0%, 92.9% and 93.2% on testing set 2, respectively. It is noted that these classifiers achieved the good performance after training with only 10% of the original training set and evaluated on the full testing set. The performance of these classifiers shows the ability of classical machine learning when the data are processed by the proper feature extraction techniques.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Machine Learning Approach for Colorectal Cancer Classification Through Learning-Based Feature Extraction

  • Le-Y. Nguyen,
  • Quoc-Hoang-Quyen Vo,
  • Minh-Khoi Ngo,
  • Cao-Bang Vu,
  • Thanh-Hai Le,
  • Thi-Thu-Hien Pham

摘要

Colorectal cancer remains a significant health challenge, necessitating the implementation of efficient methods for early identification. The application of artificial intelligence, particularly deep learning, for tissue classification holds the potential to improve diagnostic accuracy and drive breakthroughs in clinical oncology techniques. This study presents a machine learning approach using the features extracted from a pre-trained deep learning model to classify nine classes of human colorectal cancer (CRC) and normal tissue. The proposed method is trained and evaluated on the public colorectal cancer histopathological images (NCT-CRC-HE-100K and CRC-VAL-HE-7K). The training set (NCT-CRC-HE-100K) contains 100,000 histopathological images and the testing set (CRC-VAL-HE-7K) has 7180 images of nine classes. In this study, the smaller training set is created by sampling only 10,000 images of nine classes (10% of the original training set) and keeping the same ratio of these classes as the original training set. The pre-trained ResNet18 is applied to extract a vector of 512 features for each image and then the linear discriminant analysis (LDA) transformation reduces the data dimensions for the new training set and the original testing set. After that, five machine learning, i.e., Random forest (RF), Support vector machine (SVM), Light gradient boosting machine (LightGBM), Categorical gradient boosting (CatBoost) and Ensemble model (utilizing weighted average technique with the four above methods) are trained and evaluated on these sets. Furthermore, random search cross-validation technique is applied to all classifiers for finding the optimal hyperparameters. The overall accuracy of Ensemble model is highest at 89.8% and SVM, RF, LightGBM and CatBoost obtain 89.6%, 89.4%, 89.7% and 89.4% on testing set 1, respectively. Meanwhile, SVM classifier achieves the highest accuracy of 93.3% and four other classifiers (i.e., RF, LightGBM, CatBoost and Ensemble) have the accuracy of 93.0%, 93.0%, 92.9% and 93.2% on testing set 2, respectively. It is noted that these classifiers achieved the good performance after training with only 10% of the original training set and evaluated on the full testing set. The performance of these classifiers shows the ability of classical machine learning when the data are processed by the proper feature extraction techniques.