Machine learning (ML) has the potential to transform official statistics, enabling the integration of novel data sources and the development of advanced predictive capabilities based on ML algorithms. However, the rapid adoption of machine learning in this field often lacks the methodological rigor necessary to ensure statistical soundness and reliability. Based on the Total Survey Error model, we propose a framework for evaluating and addressing the sources of error in machine learning applications called the Total Machine Learning Error (TMLE) model. Using this framework, we are able to gain a grounded understanding of internal and external validity by focusing particularly on model representativeness and measurement error in future unseen data. We advocate rigorous statistical principles by situating machine learning within the broader context of official statistics. Various case studies are presented to illustrate how the TMLE framework can be applied to improve the reliability of machine learning in official statistics. A critical rethinking of machine learning’s methodology would be essential to ensure the production of high-quality, trustworthy statistics.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Leveraging Machine Learning for Official Statistics

  • Marco J. H. Puts,
  • David Salgado,
  • Piet J. H. Daas

摘要

Machine learning (ML) has the potential to transform official statistics, enabling the integration of novel data sources and the development of advanced predictive capabilities based on ML algorithms. However, the rapid adoption of machine learning in this field often lacks the methodological rigor necessary to ensure statistical soundness and reliability. Based on the Total Survey Error model, we propose a framework for evaluating and addressing the sources of error in machine learning applications called the Total Machine Learning Error (TMLE) model. Using this framework, we are able to gain a grounded understanding of internal and external validity by focusing particularly on model representativeness and measurement error in future unseen data. We advocate rigorous statistical principles by situating machine learning within the broader context of official statistics. Various case studies are presented to illustrate how the TMLE framework can be applied to improve the reliability of machine learning in official statistics. A critical rethinking of machine learning’s methodology would be essential to ensure the production of high-quality, trustworthy statistics.