Decoding Heterogeneity in Quadruple-Negative Breast Cancer: A Data-driven Clustering Approach
摘要
In a quest to decipher the complexities of Quadruple-Negative Breast Cancer (QNBC), this research harnesses advanced analytics applied to RNAseq gene expression data. Employing unsupervised clustering techniques, our rigorous methodology entails data preprocessing for enhanced interpretability, dimensionality reduction via autoencoders and Principal Component Analysis (PCA), and fine-tuning k-means clustering with internal validation indices. The analysis effectively discriminates two distinct QNBC subtypes, substantiated by high Silhouette (0.08) and Calinski-Harabasz (6.92) Scores. The profiles of these clusters are further unveiled through statistical analyses of top variant genes. Cluster 1 is typified by genes such as C9ORF57 and OR2AT4, while Cluster 2 presents distinctive genetic features, including KRTAP10-10 and ADAM3A. These data-driven clusters hold the promise of personalized assessments and interventions, contingent upon clinical validation. This study underscores the potential of integrated machine learning and statistical analysis, marking a pathway to more effective QNBC management.