Privacy-Preserving Cluster Similarity Model for Multi-user and Multi-data
摘要
This paper mainly proposes a privacy-preserving clustering similarity model for multi-user multi-group data. When different users perform clustering similarity evaluation, there is a risk of privacy leakage if the clustering results are exchanged directly. Therefore, our model advocates privacy-preserving operations on the clustering results before handing them over to a third party for similarity evaluation. The model is divided into three main stages when there are multiple sets of data: cluster ensemble, local privacy-preserving, and evaluating similarity. In the cluster ensemble stage, the method of Mean cluster ensemble is used. In the local privacy-preserving stage, the privacy-preserving method of geometric transformation and Differential privacy is used. In the assessment of similarity stage we use the commonly used clustering similarity metrics for assessment. The similarity results of clustering are given by a third party, so the use of local privacy-preserving strategies reduces the risk of privacy leakage during data transmission. In the experimental part of this paper we compare the utility of 19 commonly used similarity metrics when geometric transformation and Differential privacy are applied to cluster similarity evaluation.