Compositional data, with their constrained nature, pose unique challenges for conventional statistical approaches. To tackle these challenges, we leverage the Dirichlet distribution by presenting the contaminated unimodal Dirichlet (CUD) distribution. This variant accommodates atypical observations by offering a more flexible tail behavior. We then discuss the use of finite mixtures of CUD distributions to simultaneously handle clustering and the presence of atypical points in the data. Through simulated data analyses, we illustrate how atypical observations affect parameter estimation and data classification, showcasing the effectiveness of our approaches in addressing these issues.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Revisiting the Dirichlet Distribution for Model-Based Clustering

  • Salvatore D. Tomarchio,
  • Antonio Punzo,
  • Johannes T. Ferreira,
  • Andriette Bekker

摘要

Compositional data, with their constrained nature, pose unique challenges for conventional statistical approaches. To tackle these challenges, we leverage the Dirichlet distribution by presenting the contaminated unimodal Dirichlet (CUD) distribution. This variant accommodates atypical observations by offering a more flexible tail behavior. We then discuss the use of finite mixtures of CUD distributions to simultaneously handle clustering and the presence of atypical points in the data. Through simulated data analyses, we illustrate how atypical observations affect parameter estimation and data classification, showcasing the effectiveness of our approaches in addressing these issues.