<p>We discuss the problem of simultaneously classifying an entire group of data consisting of <i>n</i> <i>p</i>-variate observations, all in one shot, into one of the two given independent populations <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="13571_2025_392_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="20" /> </InlineMediaObject> <EquationSource Format="TEX">\(\Pi _1\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mi mathvariant="normal">Π</mi> <mn>1</mn> </msub> </math></EquationSource> </InlineEquation> or <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="13571_2025_392_Article_IEq2.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="20" /> </InlineMediaObject> <EquationSource Format="TEX">\(\Pi _2\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mi mathvariant="normal">Π</mi> <mn>2</mn> </msub> </math></EquationSource> </InlineEquation>. We assume that all observations come from the same population and thus the entire group is to be classified into either <InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="13571_2025_392_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="20" /> </InlineMediaObject> <EquationSource Format="TEX">\(\Pi _1\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mi mathvariant="normal">Π</mi> <mn>1</mn> </msub> </math></EquationSource> </InlineEquation> or <InlineEquation ID="IEq4"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="13571_2025_392_Article_IEq2.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="20" /> </InlineMediaObject> <EquationSource Format="TEX">\(\Pi _2\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mi mathvariant="normal">Π</mi> <mn>2</mn> </msub> </math></EquationSource> </InlineEquation>. In practice, these populations are represented by their respective random samples. No assumption is made about their probability distributions. Various approaches using the standard distance-based metrics along with “majority rule” exhibit very disappointing performances. Thus, we rely on eigen-structures inherent in the random samples and quantify the closeness of eigen-structure of the “to be classified sample” to the respective eigen-structures of these two random samples. Such measures are based on antieigenvalues. We also consider the associated problem of feature selection, the solution to which is also provided by devising certain suitable algorithms, which also rely on eigen-structures. The generalization to more than two populations is straightforward.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Supervised Learning and Collective Classification for a Group of Observations: An Eigen-Structure Approach

  • Huong N. Q. Tran,
  • Ravindra Khattree

摘要

We discuss the problem of simultaneously classifying an entire group of data consisting of n p-variate observations, all in one shot, into one of the two given independent populations \(\Pi _1\) Π 1 or \(\Pi _2\) Π 2 . We assume that all observations come from the same population and thus the entire group is to be classified into either \(\Pi _1\) Π 1 or \(\Pi _2\) Π 2 . In practice, these populations are represented by their respective random samples. No assumption is made about their probability distributions. Various approaches using the standard distance-based metrics along with “majority rule” exhibit very disappointing performances. Thus, we rely on eigen-structures inherent in the random samples and quantify the closeness of eigen-structure of the “to be classified sample” to the respective eigen-structures of these two random samples. Such measures are based on antieigenvalues. We also consider the associated problem of feature selection, the solution to which is also provided by devising certain suitable algorithms, which also rely on eigen-structures. The generalization to more than two populations is straightforward.