An Anonymization Technique-Based on L-Diversity for Protecting Privacy in Data Publishing
摘要
With the increasing prevalence of data collection and analysis, there is a growing concern over the privacy risks posed by personal data. Data breaches, identity theft, and other forms of cybercrime are just some of the threats from the misuse of such data. To preserve the privacy inside the data before sending them to collectors, some privacy protection approaches should be applied and data anonymization is the most widely used one in which the data is anonymized to hide the privacy. However, the deep concern of any privacy protection technique is the trade-off between data quality (i.e., the level of the difference between the original data and anonymized data) and privacy. In this research, we focus on developing an anonymization algorithm to protect the privacy of users while still ensuring that the data is processed effectively for data mining purposes (e.g., preserving data quality). The algorithm uses l-diversity as its main framework and integrates heuristics based on the tuple migration method to transform data in order to reduce the generation of new significant association rules and the loss of original ones during its execution (hence maintaining the data quality). We also conduct extensive experimentation to show the better performances of an implementation of the proposed algorithm against the state-of-the-art approaches.