Data generated by modern technologies are typically high-dimensional and noisy, that is, involve many features and large uncertainties. For example, smartphones collect information from many sensors and applications (data are high-dimensional) usually under anything but laboratory conditions (data are noisy). It turns out that such data can render statistical learning slow, unreliable, and hard to interpret. A way to gear statistical learning to high-dimensional, noisy data is regularization. Regularization complements standard pipelines with constraints on the patterns (unsupervised learning), statistical models (inferential data analyses), or prediction rules (machine learning). Regularization has a long-standing tradition in statistical learning and beyond; currently, there is a trend toward sparsity-inducing regularization, which can pinpoint and sideline irrelevant aspects of the patterns, models, or rules under consideration. This chapter discusses traditional and sparsity-inducing regularization conceptually and in concrete examples. You learn to – identify critical challenges associated with contemporary data; – implement regularization to address these challenges; – calibrate the statistical pipelines to your data.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Regularization

  • Johannes Lederer

摘要

Data generated by modern technologies are typically high-dimensional and noisy, that is, involve many features and large uncertainties. For example, smartphones collect information from many sensors and applications (data are high-dimensional) usually under anything but laboratory conditions (data are noisy). It turns out that such data can render statistical learning slow, unreliable, and hard to interpret. A way to gear statistical learning to high-dimensional, noisy data is regularization. Regularization complements standard pipelines with constraints on the patterns (unsupervised learning), statistical models (inferential data analyses), or prediction rules (machine learning). Regularization has a long-standing tradition in statistical learning and beyond; currently, there is a trend toward sparsity-inducing regularization, which can pinpoint and sideline irrelevant aspects of the patterns, models, or rules under consideration. This chapter discusses traditional and sparsity-inducing regularization conceptually and in concrete examples. You learn to – identify critical challenges associated with contemporary data; – implement regularization to address these challenges; – calibrate the statistical pipelines to your data.