Disease diagnosis is pivotal in healthcare due to its role in identifying and capturing the intricacies of various diseases. Providing appropriate patient care largely depends on accurate diagnosis. Misdiagnosis resulting from irrelevant features and noise interference may lead to inadequate or delayed treatments, unfavorable health outcomes, psychological stress, and financial loss for patients. Even as AI progresses, the venerable K-means clustering algorithms continue to be spotlighted in the domain of disease diagnosis, allowing for the revelation of insightful patterns and structures within data. However, the algorithm’s effectiveness is deeply linked to the choice of the K-means and the nature of the disease at hand. To address this challenge, this paper introduces an innovative high-dimensional K-means algorithm based on the mutual information tensors, which groups closely related features into coherent clusters and proves adept at identifying both linear and nonlinear correlations between features. By leveraging the CANDECAMP/PARAFA (CP) rank of the mutual information tensor, our method provides theoretical guarantees for the upper bound of K, which helps determine the number of clusters effectively without disrupting the structure of medical data. Empirical validation using clinical datasets has substantiated the superiority of our proposed methodology across various distance measures. Additionally, we have compared the effectiveness of different K-values between our approach and the traditional K-means method.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Approach to Determine Cluster in Tensor-Based K-Means Clustering for Disease Diagnosis

  • Liangfu Lu,
  • Shaoning Pang,
  • Joarder Kamruzzamana

摘要

Disease diagnosis is pivotal in healthcare due to its role in identifying and capturing the intricacies of various diseases. Providing appropriate patient care largely depends on accurate diagnosis. Misdiagnosis resulting from irrelevant features and noise interference may lead to inadequate or delayed treatments, unfavorable health outcomes, psychological stress, and financial loss for patients. Even as AI progresses, the venerable K-means clustering algorithms continue to be spotlighted in the domain of disease diagnosis, allowing for the revelation of insightful patterns and structures within data. However, the algorithm’s effectiveness is deeply linked to the choice of the K-means and the nature of the disease at hand. To address this challenge, this paper introduces an innovative high-dimensional K-means algorithm based on the mutual information tensors, which groups closely related features into coherent clusters and proves adept at identifying both linear and nonlinear correlations between features. By leveraging the CANDECAMP/PARAFA (CP) rank of the mutual information tensor, our method provides theoretical guarantees for the upper bound of K, which helps determine the number of clusters effectively without disrupting the structure of medical data. Empirical validation using clinical datasets has substantiated the superiority of our proposed methodology across various distance measures. Additionally, we have compared the effectiveness of different K-values between our approach and the traditional K-means method.