<p>In high-dimensional regression, the choice of regularization penalty typically forces a rigid assumption upon the underlying signal structure, dichotomizing data into strictly sparse (Lasso, <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(q=1\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi>q</mi> <mo>=</mo> <mn>1</mn> </mrow> </math></EquationSource> </InlineEquation>) or entirely dense (Ridge, <InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(q=2\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi>q</mi> <mo>=</mo> <mn>2</mn> </mrow> </math></EquationSource> </InlineEquation>) regimes. However, real-world data generating mechanisms frequently exist on a continuum between these extremes, requiring flexible geometries to handle varying degrees of sparsity and considerable multicollinearity. In this work, we propose a data-driven framework to learn the optimal regularization norm by elevating the <InlineEquation ID="IEq3"> <EquationSource Format="TEX">\(L_q\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mi>L</mi> <mi>q</mi> </msub> </math></EquationSource> </InlineEquation> exponent (<InlineEquation ID="IEq4"> <EquationSource Format="TEX">\(q \in (0, 2]\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi>q</mi> <mo>∈</mo> <mo stretchy="false">(</mo> <mn>0</mn> <mo>,</mo> <mn>2</mn> <mo stretchy="false">]</mo> </mrow> </math></EquationSource> </InlineEquation>) from a discrete choice to a strictly continuous, learnable hyper-parameter. To overcome the computational bottleneck of evaluating non-convex and non-smooth penalty landscapes, we develop a universal proximal coordinate descent solver that utilizes a safeguarded jumping threshold operator and a novel empirical Karush-Kuhn-Tucker (KKT) verification strategy. This solver is coupled with a stochastic Tree-structured Parzen Estimator (TPE) utilizing randomized internal validation splits, enabling the rapid discovery of optimal penalty geometries without over-fitting. We evaluate the framework on simulated architectures, demonstrating its dynamic adaptivity to structural sparsity, collinearity, and varying signal-to-noise ratios. Applied to four high-dimensional genomic datasets (scaling up to <InlineEquation ID="IEq5"> <EquationSource Format="TEX">\(P \approx 50,000\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi>P</mi> <mo>≈</mo> <mn>50</mn> <mo>,</mo> <mn>000</mn> </mrow> </math></EquationSource> </InlineEquation> features), our generalized adaptive bridge regression (GABR) framework successfully identifies optimal, off-grid grouping architectures (<InlineEquation ID="IEq6"> <EquationSource Format="TEX">\(q \approx 1.63\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi>q</mi> <mo>≈</mo> <mn>1.63</mn> </mrow> </math></EquationSource> </InlineEquation> to 1.80), outperforming purely sparse and purely dense alternatives. These results demonstrate that the exact regression geometry can be efficiently learned from the data, enabling a unified approach to high-dimensional inference without the computational restrictions of exhaustive discrete grid searches.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Generalized adaptive bridge regression: a unified framework for high-dimensional architecture discovery

  • Patrik Waldmann

摘要

In high-dimensional regression, the choice of regularization penalty typically forces a rigid assumption upon the underlying signal structure, dichotomizing data into strictly sparse (Lasso, \(q=1\) q = 1 ) or entirely dense (Ridge, \(q=2\) q = 2 ) regimes. However, real-world data generating mechanisms frequently exist on a continuum between these extremes, requiring flexible geometries to handle varying degrees of sparsity and considerable multicollinearity. In this work, we propose a data-driven framework to learn the optimal regularization norm by elevating the \(L_q\) L q exponent ( \(q \in (0, 2]\) q ( 0 , 2 ] ) from a discrete choice to a strictly continuous, learnable hyper-parameter. To overcome the computational bottleneck of evaluating non-convex and non-smooth penalty landscapes, we develop a universal proximal coordinate descent solver that utilizes a safeguarded jumping threshold operator and a novel empirical Karush-Kuhn-Tucker (KKT) verification strategy. This solver is coupled with a stochastic Tree-structured Parzen Estimator (TPE) utilizing randomized internal validation splits, enabling the rapid discovery of optimal penalty geometries without over-fitting. We evaluate the framework on simulated architectures, demonstrating its dynamic adaptivity to structural sparsity, collinearity, and varying signal-to-noise ratios. Applied to four high-dimensional genomic datasets (scaling up to \(P \approx 50,000\) P 50 , 000 features), our generalized adaptive bridge regression (GABR) framework successfully identifies optimal, off-grid grouping architectures ( \(q \approx 1.63\) q 1.63 to 1.80), outperforming purely sparse and purely dense alternatives. These results demonstrate that the exact regression geometry can be efficiently learned from the data, enabling a unified approach to high-dimensional inference without the computational restrictions of exhaustive discrete grid searches.