<p>Finite mixture models are ubiquitous in modern statistical modeling, and a recurring practical issue is choosing the model order. In Keribin (<i>Sankhyā Series A</i>, <b>62</b>, 49–66, 2000), the Bayesian information criterion (BIC) was proved consistent in mixtures, but under strong regularity, including high moments and high-order derivatives of the component density. We introduce the <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\nu \)</EquationSource> <EquationSource Format="MATHML"><math> <mi>ν</mi> </math></EquationSource> </InlineEquation>-BIC and <InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(\epsilon \)</EquationSource> <EquationSource Format="MATHML"><math> <mi>ϵ</mi> </math></EquationSource> </InlineEquation>-BIC, which weight the BIC penalty by negligibly small logarithmic factors immaterial in practice. This minor modification yields consistency under substantially weaker conditions, without differentiability and with mild moment assumptions, and we also give a misspecification result: when the truth lies outside the candidate family, any vanishing-penalty IC eventually selects a Kullback–Leibler optimal order among candidates. Finally, we clarify two limitations of consistent IC-based selection in mixtures: there is no universally minimal BIC-scale penalty within our sufficient conditions, and order consistency can conflict with minimax optimality in Hellinger risk. We illustrate the theory for Gaussian mixtures, non-differentiable Laplace mixtures, heavy-tailed <i>t</i>-mixtures, and mixtures of regression models.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Modifications of the BIC for order selection in finite mixture models

  • Hien Duy Nguyen,
  • TrungTin Nguyen

摘要

Finite mixture models are ubiquitous in modern statistical modeling, and a recurring practical issue is choosing the model order. In Keribin (Sankhyā Series A, 62, 49–66, 2000), the Bayesian information criterion (BIC) was proved consistent in mixtures, but under strong regularity, including high moments and high-order derivatives of the component density. We introduce the \(\nu \) ν -BIC and \(\epsilon \) ϵ -BIC, which weight the BIC penalty by negligibly small logarithmic factors immaterial in practice. This minor modification yields consistency under substantially weaker conditions, without differentiability and with mild moment assumptions, and we also give a misspecification result: when the truth lies outside the candidate family, any vanishing-penalty IC eventually selects a Kullback–Leibler optimal order among candidates. Finally, we clarify two limitations of consistent IC-based selection in mixtures: there is no universally minimal BIC-scale penalty within our sufficient conditions, and order consistency can conflict with minimax optimality in Hellinger risk. We illustrate the theory for Gaussian mixtures, non-differentiable Laplace mixtures, heavy-tailed t-mixtures, and mixtures of regression models.