<p>In the burgeoning field of artificial intelligence (AI), interaction with high-dimensional data is critical for classification problems due to noisy data points in feature variables and a lack of class separation. This paper introduces PC-NCA, a hybrid feature extraction method that links the statistical robustness of Principal Component Analysis (PCA) with the class-discriminative power of Neighborhood Component Analysis (NCA). By integrating these paradigms, the suggested method compensates for noise and redundancy and enhances class separability in high-dimensional, often non-linear data spaces. Empirical studies on 31 public datasets of varying Imbalance Ratio (1.05–18.1) across medicine, chemistry, finance, and computer security domains reveal statistically significant improvements in F1 score, G-mean, AUC, and MCC metrics, with performance improvement averaging <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(35.16\%\)</EquationSource> </InlineEquation> over the baseline and <InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(10.57\%\)</EquationSource> </InlineEquation> over traditional reduction methods. The method’s supremacy is additionally established through Wilcoxon’s signed rank statistical test (<InlineEquation ID="IEq3"> <EquationSource Format="TEX">\(p\, \text {value}&lt;0.001\)</EquationSource> </InlineEquation>) to ensure its robustness at varying magnitudes of class imbalance and feature heterogeneity. Beyond accuracy metrics, diversity analysis highlights that PC-NCA achieves <InlineEquation ID="IEq4"> <EquationSource Format="TEX">\(7.4\%\)</EquationSource> </InlineEquation> higher Neighborhood Purity (NP) and <InlineEquation ID="IEq5"> <EquationSource Format="TEX">\(62\%\)</EquationSource> </InlineEquation> lower Tomek Link Rate (TLR) compared to contemporary approaches, indicating stronger intra-class cohesion and reduced boundary ambiguity. In terms of efficiency, PC-NCA requires more runtime than PCA, LDA, and PCA-LDA (approximately <InlineEquation ID="IEq6"> <EquationSource Format="TEX">\(8\times\)</EquationSource> </InlineEquation>, <InlineEquation ID="IEq7"> <EquationSource Format="TEX">\(5\times\)</EquationSource> </InlineEquation>, and <InlineEquation ID="IEq8"> <EquationSource Format="TEX">\(3\times\)</EquationSource> </InlineEquation>, respectively) but remains consistently faster than standalone NCA while delivering comparable or superior accuracy. By integrating non-linear transformation in a denoised space, PC-NCA is an effective and adaptable alternative to the prevailing dimensionality reduction methodologies, making it a significant addition to the machine learning arsenal.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

PC-NCA: a hybrid feature extraction technique for classification in machine-learning

  • Muhammad Tanveer Islam,
  • Mahadee Al Mobin,
  • Ahsan Habib

摘要

In the burgeoning field of artificial intelligence (AI), interaction with high-dimensional data is critical for classification problems due to noisy data points in feature variables and a lack of class separation. This paper introduces PC-NCA, a hybrid feature extraction method that links the statistical robustness of Principal Component Analysis (PCA) with the class-discriminative power of Neighborhood Component Analysis (NCA). By integrating these paradigms, the suggested method compensates for noise and redundancy and enhances class separability in high-dimensional, often non-linear data spaces. Empirical studies on 31 public datasets of varying Imbalance Ratio (1.05–18.1) across medicine, chemistry, finance, and computer security domains reveal statistically significant improvements in F1 score, G-mean, AUC, and MCC metrics, with performance improvement averaging \(35.16\%\) over the baseline and \(10.57\%\) over traditional reduction methods. The method’s supremacy is additionally established through Wilcoxon’s signed rank statistical test ( \(p\, \text {value}<0.001\) ) to ensure its robustness at varying magnitudes of class imbalance and feature heterogeneity. Beyond accuracy metrics, diversity analysis highlights that PC-NCA achieves \(7.4\%\) higher Neighborhood Purity (NP) and \(62\%\) lower Tomek Link Rate (TLR) compared to contemporary approaches, indicating stronger intra-class cohesion and reduced boundary ambiguity. In terms of efficiency, PC-NCA requires more runtime than PCA, LDA, and PCA-LDA (approximately \(8\times\) , \(5\times\) , and \(3\times\) , respectively) but remains consistently faster than standalone NCA while delivering comparable or superior accuracy. By integrating non-linear transformation in a denoised space, PC-NCA is an effective and adaptable alternative to the prevailing dimensionality reduction methodologies, making it a significant addition to the machine learning arsenal.