Exploring self-supervised deep sparse autoencoders for robust feature selection in radiomics analysis
摘要
While considerable effort has been devoted to examining how variations in study protocols, acquisition settings, and annotations influence radiomics features, the role of feature selection (FS) methods largely remains unexplored. This study investigates a self-supervised deep sparse autoencoder ensemble (ensembleAE) and a novel Bayesian variant (bayesianAE) for radiomics FS in a small, class-imbalanced framework. Using a cohort of 100 prostate cancer patients under active surveillance, these models were benchmarked against eight classical FS methods. FS stability was assessed both globally and locally in a soft data-perturbation setting. While global stability measured consistency in overall feature ranking, local stability quantified the agreement in selecting top-ranked feature subsets. Among classical methods, wrappers were the least stable, whereas the filter-based Wilcoxon test (WLCX) exhibited high local stability and performance with moderate global stability. Conversely, bayesianAE demonstrated the highest global stability while matching WLCX in local stability and performance. Furthermore, although bayesianAE yielded local stability comparable to that of ensembleAE, it outperformed the latter in terms of global stability with a significantly lower computational burden (95 × faster and consumes 84% less memory). These findings position bayesianAE as a robust alternative to classical FS methods that could support the development of reproducible radiomics signatures.