<p>Spatial patterns and relationships are crucial for statistical modeling and inference across various fields. This study develops a novel approach using supervised Random Forest to compute similarity scores between locations, effectively capturing spatial dependencies of a response variable. The approach begins by enriching location coordinates, enabling Random Forest to split space into irregular shaped subspaces. The similarity score is then derived from the proportion of trees in which two locations fall in the same node for the same values of other predictors. From the resulting similarity matrix, eigen-scores and cluster labels are extracted and integrated into predictive models such as GWR, XGBoost, Random Forest, GAM, spatially and non-spatially varying coefficient (S&amp;NVC) models and Spatial Durbin Model (SDM). Simulations and two real data analyses indicate that the similarity matrix can both capture more spatial information leading to meaningful clustering results and significantly enhances the predictive performance of various models.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Supervised spatial metric learning with applications to spatial clustering and spatial model prediction

  • Xinyue Zhang,
  • Hong Gu,
  • Andrew Irwin,
  • Toby Kenney

摘要

Spatial patterns and relationships are crucial for statistical modeling and inference across various fields. This study develops a novel approach using supervised Random Forest to compute similarity scores between locations, effectively capturing spatial dependencies of a response variable. The approach begins by enriching location coordinates, enabling Random Forest to split space into irregular shaped subspaces. The similarity score is then derived from the proportion of trees in which two locations fall in the same node for the same values of other predictors. From the resulting similarity matrix, eigen-scores and cluster labels are extracted and integrated into predictive models such as GWR, XGBoost, Random Forest, GAM, spatially and non-spatially varying coefficient (S&NVC) models and Spatial Durbin Model (SDM). Simulations and two real data analyses indicate that the similarity matrix can both capture more spatial information leading to meaningful clustering results and significantly enhances the predictive performance of various models.