<p>Due to confidentiality restrictions in releasing census and survey data, such as agricultural data from the European farm structure survey (9 million records), the data are aggregated to a coarse resolution (NUTS2 administrative regions) before public release. Even when other types of census data are released as grids, grid cells may be suppressed in locations where confidentiality rules have not been respected. Here, we present a method, implemented in the <i>R</i> package <i>MRG</i>, for creating multi-resolution grids that respect restrictions while maximizing the spatial resolution at which the data are disseminated. The method can be adjusted for different restrictions, it can create the same grid structure for a set of variables, and it allows for a contextual suppression of some grid cells (i.e., suppress if all neighbors are non-confidential, merge if several others are also confidential) if this results in a generally higher information content, a combination of features that has not previously been available. The method is exemplified with a synthetic data set.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A flexible approach for statistical disclosure control in geospatial data

  • Jon Olav Skøien,
  • Nicolas Lampach,
  • Helena Ramos,
  • Rudolf Seljak,
  • Renate Koeble,
  • Linda See,
  • Marijn van der Velde

摘要

Due to confidentiality restrictions in releasing census and survey data, such as agricultural data from the European farm structure survey (9 million records), the data are aggregated to a coarse resolution (NUTS2 administrative regions) before public release. Even when other types of census data are released as grids, grid cells may be suppressed in locations where confidentiality rules have not been respected. Here, we present a method, implemented in the R package MRG, for creating multi-resolution grids that respect restrictions while maximizing the spatial resolution at which the data are disseminated. The method can be adjusted for different restrictions, it can create the same grid structure for a set of variables, and it allows for a contextual suppression of some grid cells (i.e., suppress if all neighbors are non-confidential, merge if several others are also confidential) if this results in a generally higher information content, a combination of features that has not previously been available. The method is exemplified with a synthetic data set.