Hierarchical Clustering with an Ensemble of Principle Component Trees for Interpretable Patient Stratification
摘要
Patient stratification plays a crucial role in personalized medicine by identifying distinct subgroups of patients based on their molecular and/or clinical characteristics. However, many unsupervised machine learning-based stratification techniques fail to identify the essential biomarker traits associated with each patient group. In this paper, we present a novel approach for interpretable patient stratification using hierarchical ensemble clustering. Our method leverages feature sampling in conjunction with principal component analysis (PCA) to capture the most significant patterns and contributing biomarkers. We demonstrate the effectiveness of our approach using machine learning benchmark datasets and real-world data from The Cancer Genome Atlas (TCGA), showcasing the improved interpretability of the detected patient clusters.