The amount of data generated is growing exponentially every year and such data can be collected and then used to discover valuable information through data mining processes. In fact, data may contain sensitive information; thus, it is urgent to protect that sensitive data before sharing the data to the collectors. This paper proposes an efficient algorithm to transform the original data to achieve k-anonymity, a well-known privacy model, while preserving significant information in the original data. This guarantees that the data mining process based on association rule mining after that can mine valuable information as in the original data. The key idea of the proposed algorithm is to group data records that are the same in the quasi-identifiers and then perform efficient tuple migrations to improve the data quality while still obtaining k-anonymity. An extensive experimentation reports the better performances of an implementation of our algorithm with respect to the state-of-the-art approaches in a variety of settings of data quality.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A K-Anonymity Technique for Privacy Preserving Data Publishing

  • Anh Tuan Truong

摘要

The amount of data generated is growing exponentially every year and such data can be collected and then used to discover valuable information through data mining processes. In fact, data may contain sensitive information; thus, it is urgent to protect that sensitive data before sharing the data to the collectors. This paper proposes an efficient algorithm to transform the original data to achieve k-anonymity, a well-known privacy model, while preserving significant information in the original data. This guarantees that the data mining process based on association rule mining after that can mine valuable information as in the original data. The key idea of the proposed algorithm is to group data records that are the same in the quasi-identifiers and then perform efficient tuple migrations to improve the data quality while still obtaining k-anonymity. An extensive experimentation reports the better performances of an implementation of our algorithm with respect to the state-of-the-art approaches in a variety of settings of data quality.