This paper deals with the optimization of classification trees. We discuss the case of small training samples and large test samples so that the model accuracy on the test set is adequate for the characterization of the model’s predictive power. The idea of our approach is to repeatedly generate models on small subsamples of the overall learning sample and identify that model which is best on the corresponding test sample. In our application, not the number of misclassifications of the individual observations is relevant for the model accuracy, but the number of misclassified sequences of observations of a specified minimum length. This leads to an innovative optimization criterion. Our application belongs to Music Data Analysis.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Optimal Classification Trees: Sequential Analysis

  • Claus Weihs

摘要

This paper deals with the optimization of classification trees. We discuss the case of small training samples and large test samples so that the model accuracy on the test set is adequate for the characterization of the model’s predictive power. The idea of our approach is to repeatedly generate models on small subsamples of the overall learning sample and identify that model which is best on the corresponding test sample. In our application, not the number of misclassifications of the individual observations is relevant for the model accuracy, but the number of misclassified sequences of observations of a specified minimum length. This leads to an innovative optimization criterion. Our application belongs to Music Data Analysis.