<p>Spatial averaging has become an important tool in spatial data analysis, achieved by assigning different weights to various spatial observations. Missing data is common in many spatial datasets. In this article, we propose a form of spatial average statistics for both random and non-random missing data and present the optimal sampling weights by minimizing bias and variance, respectively. The optimal weights depend on the missing mechanism. Compared to traditional logistic regression models, we consider machine learning techniques, such as Classification and Regression Trees and Random Forests, as alternative candidates for modeling the missing mechanism. We employ two simulation designs in this study: a general data generation mechanism and a simulation that mimics real data by maintaining the structure and confounding factors consistent with the real dataset. The results show that the optimal averaging strategy proposed in this paper can closely approximate the true value, regardless of the missing mechanism. Finally, we apply the proposed method to the Tara Oceans data and estimate the optimal weights for the spatial average of the relative abundance of Ammonia-Oxidizing Archaea.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

The optimal spatial averaging method for random/non-random missing data via super learner and its application to Tara Oceans data

  • Aixian Chen,
  • Xiaoyun Huang,
  • Xia Cui

摘要

Spatial averaging has become an important tool in spatial data analysis, achieved by assigning different weights to various spatial observations. Missing data is common in many spatial datasets. In this article, we propose a form of spatial average statistics for both random and non-random missing data and present the optimal sampling weights by minimizing bias and variance, respectively. The optimal weights depend on the missing mechanism. Compared to traditional logistic regression models, we consider machine learning techniques, such as Classification and Regression Trees and Random Forests, as alternative candidates for modeling the missing mechanism. We employ two simulation designs in this study: a general data generation mechanism and a simulation that mimics real data by maintaining the structure and confounding factors consistent with the real dataset. The results show that the optimal averaging strategy proposed in this paper can closely approximate the true value, regardless of the missing mechanism. Finally, we apply the proposed method to the Tara Oceans data and estimate the optimal weights for the spatial average of the relative abundance of Ammonia-Oxidizing Archaea.