<p>High-dimensional regression is often complicated by outliers and missing data, which can undermine both accuracy and interpretability in variable selection. We propose a unified robust penalized regression framework that integrates Huber loss to mitigate the impact of outliers and employs an Expectation-Maximization (EM) algorithm to impute missing covariates. This framework enables principled and simultaneous handling of contamination and missingness. Within this framework, we introduce two <i>new</i> strategies. The first is the HYBRID Penalty Method, a novel nonconvex penalty that synthesizes gradient information from SCAD, MCP, and LOG to form an adaptive penalty function with reduced bias and improved stability. The second is the ENSEMBLE Selection Method, a new stability-enhancing mechanism that aggregates multiple penalized selections through a majority-voting rule, thereby increasing reproducibility in high-dimensional settings. Both methods are implemented via an efficient coordinate descent algorithm with local linear approximation, ensuring computational scalability. Extensive simulation studies and real-data applications show that HYBRID and ENSEMBLE substantially outperform existing robust and imputation-based techniques in terms of support recovery, prediction accuracy, and resilience to data irregularities. These results highlight the novelty and practical value of our proposed framework for reliable variable selection under missing data and contamination.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Introducing HYBRID and ENSEMBLE: novel nonconvex penalization strategies for robust variable selection under missing data

  • Lingling Zhang

摘要

High-dimensional regression is often complicated by outliers and missing data, which can undermine both accuracy and interpretability in variable selection. We propose a unified robust penalized regression framework that integrates Huber loss to mitigate the impact of outliers and employs an Expectation-Maximization (EM) algorithm to impute missing covariates. This framework enables principled and simultaneous handling of contamination and missingness. Within this framework, we introduce two new strategies. The first is the HYBRID Penalty Method, a novel nonconvex penalty that synthesizes gradient information from SCAD, MCP, and LOG to form an adaptive penalty function with reduced bias and improved stability. The second is the ENSEMBLE Selection Method, a new stability-enhancing mechanism that aggregates multiple penalized selections through a majority-voting rule, thereby increasing reproducibility in high-dimensional settings. Both methods are implemented via an efficient coordinate descent algorithm with local linear approximation, ensuring computational scalability. Extensive simulation studies and real-data applications show that HYBRID and ENSEMBLE substantially outperform existing robust and imputation-based techniques in terms of support recovery, prediction accuracy, and resilience to data irregularities. These results highlight the novelty and practical value of our proposed framework for reliable variable selection under missing data and contamination.