The identification of significant input variables is crucial for the quality of machine learning models, since the models can only be as good as the representative quality of the data provided. In this chapter, different filter methods are compared as a subset of data analysis techniques for feature selection. Using real data from the power engineering field, the strengths of data analysis as a preprocessing step are demonstrated and the effectiveness of different filter methods is compared. The data analysis techniques not only increase the performance of machine learning models, but also their plausibility by identifying important input variables and reducing model complexity, resulting in more trustworthy and reliable models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Increasing the Performance and Plausibility of Machine Learning via Data Analysis Techniques

  • Silas Aaron Selzer,
  • Fabian Bauer,
  • Peter Bretschneider

摘要

The identification of significant input variables is crucial for the quality of machine learning models, since the models can only be as good as the representative quality of the data provided. In this chapter, different filter methods are compared as a subset of data analysis techniques for feature selection. Using real data from the power engineering field, the strengths of data analysis as a preprocessing step are demonstrated and the effectiveness of different filter methods is compared. The data analysis techniques not only increase the performance of machine learning models, but also their plausibility by identifying important input variables and reducing model complexity, resulting in more trustworthy and reliable models.