<p>In this paper, we propose a novel method for clustering functional data that significantly extends the distribution scope of functional data. Most existing methods are developed for a specific distribution structure of functional data and, thus, provide good results only when the distribution assumptions are met. For example, mean-based methods perform poorly when the error distribution is asymmetric. In contrast, methods based on asymmetric measures perform relatively poorly compared to mean-based methods in the presence of symmetric errors. In addition, most methods based on the basis expansion approach use only one curve representing the centrality of data as a clustering input, limiting their ability to handle complex distributional structures. To address this limitation, we consider multiple types of curves, including mean and quantile curves, and use their functional principal component scores as input variables for clustering, which can better reflect the distribution of data. Furthermore, we apply the concept of sparse clustering to multiple types of curves, resulting in good clustering performance if at least one curve can divide data into several subgroups well. Results from numerical experiments, including real data analysis, empirically verify that the proposed approach consistently achieves superior clustering performance across various distributional settings.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

RSFclust: Robust sparse clustering of functional data using quantile curves

  • Chihoon Lee,
  • Hee-Seok Oh,
  • Joonpyo Kim

摘要

In this paper, we propose a novel method for clustering functional data that significantly extends the distribution scope of functional data. Most existing methods are developed for a specific distribution structure of functional data and, thus, provide good results only when the distribution assumptions are met. For example, mean-based methods perform poorly when the error distribution is asymmetric. In contrast, methods based on asymmetric measures perform relatively poorly compared to mean-based methods in the presence of symmetric errors. In addition, most methods based on the basis expansion approach use only one curve representing the centrality of data as a clustering input, limiting their ability to handle complex distributional structures. To address this limitation, we consider multiple types of curves, including mean and quantile curves, and use their functional principal component scores as input variables for clustering, which can better reflect the distribution of data. Furthermore, we apply the concept of sparse clustering to multiple types of curves, resulting in good clustering performance if at least one curve can divide data into several subgroups well. Results from numerical experiments, including real data analysis, empirically verify that the proposed approach consistently achieves superior clustering performance across various distributional settings.