Revisiting the Dirichlet Distribution for Model-Based Clustering
摘要
Compositional data, with their constrained nature, pose unique challenges for conventional statistical approaches. To tackle these challenges, we leverage the Dirichlet distribution by presenting the contaminated unimodal Dirichlet (CUD) distribution. This variant accommodates atypical observations by offering a more flexible tail behavior. We then discuss the use of finite mixtures of CUD distributions to simultaneously handle clustering and the presence of atypical points in the data. Through simulated data analyses, we illustrate how atypical observations affect parameter estimation and data classification, showcasing the effectiveness of our approaches in addressing these issues.