Abstract <p>The article presents a new methodology for solving problems of self-supervised and semisupervised learning using deep Gaussian mixture models for the case of small datasets for which only a limited set of training data is typically available. On the basis of a mathematical model implemented in the layers of a neural network, a method is proposed for forming pseudolabels of data, which are subsequently used to train a classifier or regressor, for which various machine learning methods and neural network architectures are used in the research. The deep Gaussian mixture models allows us to form an enriched feature space for describing data, expanding the approach of probability-informed machine learning developed by the authors. To test the method, 19 open tabular datasets from the University of California Irvine repository, designed to solve various machine learning problems, as well as time series describing the change in the radiation flux of type Ia supernovae (binary stars), were used. It has been established that, in the case where the labeled data constitutes less than half of the set, the methodology proposed in the article allows using more complex and deep architectures (for example, transformer ones) for processing and obtaining results with better accuracy with their help. With a large proportion of labeled data, the use of deep Gaussian mixture models allows us to increase the accuracy of simpler algorithms to the level of more complex ones (in terms of the number of trainable parameters and architecture) without using additional observations. An increase in the accuracy of solving supervised learning problems using pseudolabels generated by deep Gaussian mixture models is demonstrated, compared to traditionally used labeling approaches, such as classical Gaussian mixture models and k-means. Thus, the average increase in the F1-score in the classification task amounted to as much as 19.26%, and the average relative decrease in the mean square error in the regression problem reached 36.74% on test datasets.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Small Sample Classification and Regression for Time Series and Tabular Data Based on Deep Gaussian Mixture Models

  • A. K. Gorshenin,
  • A. M. Dostovalova

摘要

Abstract

The article presents a new methodology for solving problems of self-supervised and semisupervised learning using deep Gaussian mixture models for the case of small datasets for which only a limited set of training data is typically available. On the basis of a mathematical model implemented in the layers of a neural network, a method is proposed for forming pseudolabels of data, which are subsequently used to train a classifier or regressor, for which various machine learning methods and neural network architectures are used in the research. The deep Gaussian mixture models allows us to form an enriched feature space for describing data, expanding the approach of probability-informed machine learning developed by the authors. To test the method, 19 open tabular datasets from the University of California Irvine repository, designed to solve various machine learning problems, as well as time series describing the change in the radiation flux of type Ia supernovae (binary stars), were used. It has been established that, in the case where the labeled data constitutes less than half of the set, the methodology proposed in the article allows using more complex and deep architectures (for example, transformer ones) for processing and obtaining results with better accuracy with their help. With a large proportion of labeled data, the use of deep Gaussian mixture models allows us to increase the accuracy of simpler algorithms to the level of more complex ones (in terms of the number of trainable parameters and architecture) without using additional observations. An increase in the accuracy of solving supervised learning problems using pseudolabels generated by deep Gaussian mixture models is demonstrated, compared to traditionally used labeling approaches, such as classical Gaussian mixture models and k-means. Thus, the average increase in the F1-score in the classification task amounted to as much as 19.26%, and the average relative decrease in the mean square error in the regression problem reached 36.74% on test datasets.