Distributed sparse learning with high-dimensional and non-randomly stored data
摘要
Distributed sparse learning (DSL) serves as a pivotal framework for the collaborative analysis of high-dimensional data collected across multiple machines. However, in many practical settings, data are distributed non-randomly, leading to heterogeneous local datasets. This heterogeneity presents substantial challenges to many conventional DSL methods, which are designed for randomly distributed data. To address this issue, in this paper, we study sparse learning with non-randomly distributed data. We propose an efficient and adaptive distributed sparse learning (ADSL) approach, which is designed to handle high-dimensional data robustly under both random and non-random data storage mechanisms. The proposed method achieves a favorable balance between statistical efficiency and communication efficiency, while remaining computationally scalable. We further establish theoretical guarantees for both parameter estimation consistency and model selection consistency, ensuring the reliability of ADSL in both randomly and non-randomly distributed settings. Extensive numerical experiments and real data analysis are conducted to corroborate the theoretical findings, demonstrating that the proposed method outperforms existing approaches in terms of estimation accuracy, computational efficiency, and communication cost.