Domain-adaptive high-dimensional variable selection via transfer learning and stability fusion
摘要
High-dimensional variable selection under domain adaptation poses a critical challenge in modern machine learning. Traditional selection methods often fail to generalize when the training (source) and testing (target) distributions differ, particularly in scenarios where the target domain lacks labeled data. In this paper, we propose a novel method that integrates transferable representation learning, model-based sensitivity analysis, and statistical stability selection to identify variables that are both predictive and transferable. Specifically, we design a three-stage framework that (i) learns domain-invariant embeddings using distribution alignment, (ii) computes variable importance through gradient attribution and subsample-based stability analysis, and (iii) fuses these perspectives into a robust variable selection strategy. We further formulate the entire process within a rigorous mathematical framework and provide theoretical guarantees, including a generalization bound for target risk, stability-based false selection control, and information retention of selected variables. Experimental results on both synthetic and real-world datasets demonstrate that our approach outperforms existing methods in terms of both predictive accuracy and variable interpretability across domains.