Big data subsampling encompasses a set of techniques aimed at selecting the most informative subset from an initial dataset, with the purpose of conducting inferential analyses on said subset, thereby reducing the computational power required to perform such analyses. We present a literature review on the topic with an orientation specifically designed for practitioners. As such we will introduce some metrics used to analyze the sample of articles in the review and discuss them, then briefly present one of the most commonly used subsampling methods, highlighting its advantages and limitations.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Big Data Subsampling: A Review

  • Rosa Arboretti,
  • Marta Disegna,
  • Alberto Molena

摘要

Big data subsampling encompasses a set of techniques aimed at selecting the most informative subset from an initial dataset, with the purpose of conducting inferential analyses on said subset, thereby reducing the computational power required to perform such analyses. We present a literature review on the topic with an orientation specifically designed for practitioners. As such we will introduce some metrics used to analyze the sample of articles in the review and discuss them, then briefly present one of the most commonly used subsampling methods, highlighting its advantages and limitations.