As the application of machine learning methods continues to gain relevance in the field of official statistics, the accurate estimation of model performance becomes increasingly crucial. Resampling methods, such as cross-validation and bootstrap, offer powerful tools for estimating the generalization error of predictive models. This chapter explores the concept of generalization error and delves into the intricate challenges associated with conducting inference on it. Building upon the foundational insights, we provide a comprehensive survey of existing methodologies in the literature, offering assessments and recommendations for both uncertainty estimation via confidence intervals and the derivation of point estimates within the context of nonstandard data structures, in particular, non-i.i.d. data.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Challenges in Resampling-Based Performance Estimation

  • Hannah Schulz-Kümpel,
  • Anne-Laure Boulesteix,
  • Sebastian Fischer,
  • Roman Hornung

摘要

As the application of machine learning methods continues to gain relevance in the field of official statistics, the accurate estimation of model performance becomes increasingly crucial. Resampling methods, such as cross-validation and bootstrap, offer powerful tools for estimating the generalization error of predictive models. This chapter explores the concept of generalization error and delves into the intricate challenges associated with conducting inference on it. Building upon the foundational insights, we provide a comprehensive survey of existing methodologies in the literature, offering assessments and recommendations for both uncertainty estimation via confidence intervals and the derivation of point estimates within the context of nonstandard data structures, in particular, non-i.i.d. data.