Information processing and decision-making are inseparable from the role of data classification and prediction. Therefore, this paper proposes an integrated learning strategy to solve this problem, combining the advantages of random forest (RF) and gradient lifting machine (GBM) to improve the accuracy and efficiency of data classification. In the method part, the whole process from data preprocessing and feature selection to RF and GBM algorithm implementation is discussed. In the results and discussion section, the effectiveness of the integrated model is demonstrated by several performance indicators. The results show that the integrated model is superior to the separate RF and GBM models in prediction time, CPU utilization and throughput. Specifically, the prediction time of the integrated model is much less than that of the other two models, the average CPU utilization rate is reduced to 34.2%, and the throughput rate is mostly kept above 1000 KB/s. These results show that the integrated model has high efficiency and excellent performance in processing large-scale data. Although set models are complicated in construction and interpretation, they have high accuracy and generalization ability in data classification tasks, so they are an ideal choice when the problems in classification become more complicated.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Efficient Data Classification and Prediction Using Random Forest (RF) and Gradient Boosting Machine (GBM)

  • Junliang Du,
  • Xiaoyi Wang,
  • Junpeng Chen,
  • Ziyan Zhao,
  • Yang Zheng

摘要

Information processing and decision-making are inseparable from the role of data classification and prediction. Therefore, this paper proposes an integrated learning strategy to solve this problem, combining the advantages of random forest (RF) and gradient lifting machine (GBM) to improve the accuracy and efficiency of data classification. In the method part, the whole process from data preprocessing and feature selection to RF and GBM algorithm implementation is discussed. In the results and discussion section, the effectiveness of the integrated model is demonstrated by several performance indicators. The results show that the integrated model is superior to the separate RF and GBM models in prediction time, CPU utilization and throughput. Specifically, the prediction time of the integrated model is much less than that of the other two models, the average CPU utilization rate is reduced to 34.2%, and the throughput rate is mostly kept above 1000 KB/s. These results show that the integrated model has high efficiency and excellent performance in processing large-scale data. Although set models are complicated in construction and interpretation, they have high accuracy and generalization ability in data classification tasks, so they are an ideal choice when the problems in classification become more complicated.