Deriving Biologically Relevant Rules in Breast Cancer Subtypes Using FP-Growth Algorithm
摘要
This study presents a comprehensive framework for data mining to identify significant biomarkers and gene interactions in various breast cancer subtypes using gene expression data. The methodology employs robust multi-model techniques, including feature selection, fuzzy logic-based discretization, and FP-Growth association rule mining, to uncover interpretable gene expression patterns. Utilizing two high-dimensional datasets—focusing on proteomics (D1) and transcriptomics (D2)—this approach allows for both general and subtype-specific analysis of breast cancer. The key findings highlighted biologically relevant genes, such as EGFR, TP53, ERBB2, and VEGFA, alongside strong association rules with high support and confidence. Additionally, network analysis revealed distinct connectivity patterns among the subtypes, providing insights into potential biomarkers and therapeutic targets. This study demonstrates the effectiveness of integrating multi-model feature selection with rule-based mining to identify meaningful gene associations pertinent to breast cancer diagnosis and treatment.