Enhancing Breast Cancer Diagnosis Through Adaptive k-means Clustering and Feature Selection
摘要
This study presents an approach for enhancing breast cancer diagnosis using the Wisconsin Breast Cancer dataset. The dataset, comprising 699 patient records, contains detailed nuclear features from fine needle aspirations of breast tissues. An adaptive k-means clustering algorithm (AKMA) was adopted, integrating advanced feature selection techniques and parameter settings. In the training of the AKMA on the Wisconsin Breast Cancer dataset, the algorithm initially selects centroids based on heuristic seeds, optimizing for robust starting points. The process involves iteratively reallocating data points to clusters based on a composite of Euclidean and Manhattan distances, recalculating centroids until stable. The algorithm's performance is evaluated using positive predictive value (PPV) as a metric, indicating classification accuracy. Results demonstrate the effectiveness of the approach, achieving a remarkable PPV of 90% and accuracy of 92% considering different epochs and iterations on the Wisconsin Diagnostic Breast Cancer dataset. The PPV shows an upward trend with increasing training epochs, emphasizing the model's robustness. This approach offers a promising solution for improving breast cancer diagnosis accuracy, especially for small tumors within dense breast tissue. Further research and clinical validation are warranted to fully realize its potential impact on breast cancer care.