Similarity Measures for Datasets with Binary Variables and Their Application in Agglomerative Cluster Analysis
摘要
This paper focuses on similarity measures applicable in agglomerative clustering analysis of datasets with binary variables and their impact on resulting clustering solutions. Specifically, it analyzes 65 measures for binary data. The analysis of the influence of selecting a measure on the outputs of clustering analysis is conducted through a simulation study. The mutual similarity of clustering solutions and the quality of clustering solutions resulting from individual measures are evaluated, using both internal and external evaluation criteria, while two clustering methods are used in the experiment. Finally, two groups of measures leading to almost identical clustering solutions were identified. This way the impact of the measure on clustering results is examined and quantified, which an area that has not been sufficiently explored until now.