Synthetic data: revisiting the privacy-utility trade-off
摘要
Synthetic data is regarded as a better privacy-preserving alternative to traditionally sanitized data across various applications. However, a recent article challenges this notion, stating that synthetic data does not provide a better trade-off between privacy and utility than traditional anonymization techniques, and that it leads to unpredictable utility loss and highly unpredictable privacy gain. The article also claims to have identified a breach in the differential privacy guarantees provided by PATE-GAN and PrivBayes. Our analysis indicates that when evaluations are conducted in highly specialized and constrained environments, the generalizability of the findings is limited. Moreover, we observed that when key preconditions related to data distributions are not met in experiments, it may lead to spurious observation of violations of the differential privacy guarantee. Subsequently, we performed a comparative privacy-utility analysis using more generalized and unconstrained settings. Our results indicate that although not all synthetic data generation techniques outperformed k-anonymization across every dataset, for each dataset, at least one generator yielded a more favorable privacy-utility trade-off than the k-anonymization method, thereby reaffirming earlier conclusions.