<p>Subsampling algorithms for various parametric regression models with massive data have been extensively investigated in recent years. However, all existing studies on subsampling heavily rely on clean massive data. In practical applications, the observed covariates may suffer from inaccuracies due to measurement errors. To address the challenge of large datasets with measurement errors, this study explores two subsampling algorithms based on the corrected likelihood approach: the optimal subsampling algorithm utilizing inverse probability weighting and the perturbation subsampling algorithm employing random weighting with a perfectly known distribution. Theoretical properties for both algorithms are provided. Numerical simulations and two real-data examples demonstrate the effectiveness of these proposed methods compared to other existing algorithms.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Subsampling for Big Data Linear Models with Measurement Errors

  • Jiangshan Ju,
  • Min-Qian Liu,
  • Mingqiu Wang,
  • Shengli Zhao

摘要

Subsampling algorithms for various parametric regression models with massive data have been extensively investigated in recent years. However, all existing studies on subsampling heavily rely on clean massive data. In practical applications, the observed covariates may suffer from inaccuracies due to measurement errors. To address the challenge of large datasets with measurement errors, this study explores two subsampling algorithms based on the corrected likelihood approach: the optimal subsampling algorithm utilizing inverse probability weighting and the perturbation subsampling algorithm employing random weighting with a perfectly known distribution. Theoretical properties for both algorithms are provided. Numerical simulations and two real-data examples demonstrate the effectiveness of these proposed methods compared to other existing algorithms.