Abstract <p>To cluster, classify and represent are three fundamental objectives of learning from high-dimensional data with intrinsic structure. To this end, three interpretable approaches, i.e., segmentation (clustering) via the minimum lossy coding length criterion, classification via the minimum incremental coding length criterion and representation via the maximal coding rate reduction criterion, are introduced. These algorithms were derived based on the lossy data coding and compression framework from the principle of rate distortion in information theory. They are particularly suitable for dealing with finite-sample data (allowed to be sparse or almost degenerate) of mixed Gaussian distributions or subspaces. The theoretical value and attractive features of these methods are summarized by comparison with other learning methods or evaluation criteria. This essay aims to provide a theoretical guide to researchers (especially engineers) interested in understanding “white-box” machine (deep) learning methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

On Interpretable Approaches to Cluster, Classify, and Represent Multisubspace Data via Minimum Lossy Coding Length Based on Rate-Distortion Theory

  • Kai-liang Lu,
  • Avraham Chapman

摘要

Abstract

To cluster, classify and represent are three fundamental objectives of learning from high-dimensional data with intrinsic structure. To this end, three interpretable approaches, i.e., segmentation (clustering) via the minimum lossy coding length criterion, classification via the minimum incremental coding length criterion and representation via the maximal coding rate reduction criterion, are introduced. These algorithms were derived based on the lossy data coding and compression framework from the principle of rate distortion in information theory. They are particularly suitable for dealing with finite-sample data (allowed to be sparse or almost degenerate) of mixed Gaussian distributions or subspaces. The theoretical value and attractive features of these methods are summarized by comparison with other learning methods or evaluation criteria. This essay aims to provide a theoretical guide to researchers (especially engineers) interested in understanding “white-box” machine (deep) learning methods.