PC-NCA: a hybrid feature extraction technique for classification in machine-learning
摘要
In the burgeoning field of artificial intelligence (AI), interaction with high-dimensional data is critical for classification problems due to noisy data points in feature variables and a lack of class separation. This paper introduces PC-NCA, a hybrid feature extraction method that links the statistical robustness of Principal Component Analysis (PCA) with the class-discriminative power of Neighborhood Component Analysis (NCA). By integrating these paradigms, the suggested method compensates for noise and redundancy and enhances class separability in high-dimensional, often non-linear data spaces. Empirical studies on 31 public datasets of varying Imbalance Ratio (1.05–18.1) across medicine, chemistry, finance, and computer security domains reveal statistically significant improvements in F1 score, G-mean, AUC, and MCC metrics, with performance improvement averaging