This chapter presents an innovative automated machine learning pipeline for statistical matching, aimed at improving the efficiency and quality of data processing in official statistics. The pipeline simplifies the entire workflow, from data ingestion and preprocessing to modeling and reporting, making it easier to handle complex data sets. This pipeline leverages modern machine learning algorithms to effectively capture nonlinear relationships but remain rarely used during matching in official statistics. A simulation study using data from the European Social Survey demonstrates that the pipeline leads to moderate accuracy in predicting values and relationships. Additionally, a replication study of a real-world project highlights the pipeline’s efficiency, significantly reducing development time while achieving results comparable to conventional approaches. This work emphasizes the potential of automated techniques in statistical matching, enabling better resource allocation and quicker responses to evolving data needs in official statistics. By making these advanced methods more accessible, there is potential to enhance the quality and timeliness of statistical information for decision-makers and the public alike.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Automated Machine Learning Pipeline for Statistical Matching

  • Theresa Küntzler

摘要

This chapter presents an innovative automated machine learning pipeline for statistical matching, aimed at improving the efficiency and quality of data processing in official statistics. The pipeline simplifies the entire workflow, from data ingestion and preprocessing to modeling and reporting, making it easier to handle complex data sets. This pipeline leverages modern machine learning algorithms to effectively capture nonlinear relationships but remain rarely used during matching in official statistics. A simulation study using data from the European Social Survey demonstrates that the pipeline leads to moderate accuracy in predicting values and relationships. Additionally, a replication study of a real-world project highlights the pipeline’s efficiency, significantly reducing development time while achieving results comparable to conventional approaches. This work emphasizes the potential of automated techniques in statistical matching, enabling better resource allocation and quicker responses to evolving data needs in official statistics. By making these advanced methods more accessible, there is potential to enhance the quality and timeliness of statistical information for decision-makers and the public alike.