<p>In the contemporary era of technological advancements, the accessibility and sharing of information have increased significantly. The vast amounts of data, commonly known as big data, present an opportunity to extract hidden knowledge through various clustering algorithms. A wide variety of clustering algorithms has been developed to cluster big data into groups. However, medical records usually contain a large amount of information to be processed and this requires a lot of computing power, memory size and processing time. Hence, an efficient clustering algorithm is needed. This study explores the application of data clustering on cancer data to extract hidden knowledge from clustering plots by implementing the balanced iterative reducing and clustering using hierarchies (BIRCH) algorithm. The surveillance, epidemiology, and end results program (SEER), initiated by the National Cancer Institute, collected a substantial cancer patients’ dataset which has been used in this study to apply the BIRCH algorithm. To improve the BIRCH algorithm, this study concentrates on the distance measuring technique during centroid calculation. Comparing Euclidean and Manhattan distances, we conducted experiments using breast, leukemia, and stomach cancer data from SEER dataset. Injecting a 4-dimensional cancer patient sample data (“Age Group”, “Year of Diagnosis”, “Cancer Type”, and “Cancer Degree”) into the BIRCH algorithm, we assessed its performance with both distance techniques across various clusters (2–9). The evaluation considered qualitative aspects of clustering plots and quantitative aspects of implementation time. Our findings indicate that Manhattan distance outperformed Euclidean distance across all assigned cluster numbers, establishing its superiority in the BIRCH algorithm’s implementation.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Improving BIRCH hierarchical clustering algorithms for enhanced partitioning of medical data

  • Faiz Alshrouf,
  • Shadi Masadeh,
  • Dimah Al-Fraihat,
  • Raed Al-Nsoor

摘要

In the contemporary era of technological advancements, the accessibility and sharing of information have increased significantly. The vast amounts of data, commonly known as big data, present an opportunity to extract hidden knowledge through various clustering algorithms. A wide variety of clustering algorithms has been developed to cluster big data into groups. However, medical records usually contain a large amount of information to be processed and this requires a lot of computing power, memory size and processing time. Hence, an efficient clustering algorithm is needed. This study explores the application of data clustering on cancer data to extract hidden knowledge from clustering plots by implementing the balanced iterative reducing and clustering using hierarchies (BIRCH) algorithm. The surveillance, epidemiology, and end results program (SEER), initiated by the National Cancer Institute, collected a substantial cancer patients’ dataset which has been used in this study to apply the BIRCH algorithm. To improve the BIRCH algorithm, this study concentrates on the distance measuring technique during centroid calculation. Comparing Euclidean and Manhattan distances, we conducted experiments using breast, leukemia, and stomach cancer data from SEER dataset. Injecting a 4-dimensional cancer patient sample data (“Age Group”, “Year of Diagnosis”, “Cancer Type”, and “Cancer Degree”) into the BIRCH algorithm, we assessed its performance with both distance techniques across various clusters (2–9). The evaluation considered qualitative aspects of clustering plots and quantitative aspects of implementation time. Our findings indicate that Manhattan distance outperformed Euclidean distance across all assigned cluster numbers, establishing its superiority in the BIRCH algorithm’s implementation.