Data anonymization has become an essential preprocessing step in many analytics, especially the ones that rely on machine learning. Retrieving external data sources can significantly improve the quality of machine learning models but the sensitive nature often hinders data exchange. On the other hand, anonymization tactics can substantially decrease their utility resulting in inferior machine learning models. This work assesses the quality of anonymized datasets for multi-class classification purposes. A hybrid anonymization pipeline is employed, combining masking and sampling. Deliberate sampling – achieved by balancing the target attribute – not only offers strong and quantifiable privacy guarantees but also improves the utility of released datasets. Our findings are supported by experiments that are executed on three distinct datasets, demonstrating the effectiveness of the approach.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Advanced Strategies for Privacy Preserving Data Publishing to Improve Multi-class Classification

  • Tibo Laperre,
  • Jenno Verdonck,
  • Kevin De Boeck,
  • Michiel Willocx,
  • Vincent Naessens

摘要

Data anonymization has become an essential preprocessing step in many analytics, especially the ones that rely on machine learning. Retrieving external data sources can significantly improve the quality of machine learning models but the sensitive nature often hinders data exchange. On the other hand, anonymization tactics can substantially decrease their utility resulting in inferior machine learning models. This work assesses the quality of anonymized datasets for multi-class classification purposes. A hybrid anonymization pipeline is employed, combining masking and sampling. Deliberate sampling – achieved by balancing the target attribute – not only offers strong and quantifiable privacy guarantees but also improves the utility of released datasets. Our findings are supported by experiments that are executed on three distinct datasets, demonstrating the effectiveness of the approach.