Synthetic Data for Privacy Preservation in Distributed Data Analysis Systems
摘要
Over the years, data anonymization and federated learning have been proposed to address the challenges of data privacy in distributed systems. Unfortunately, data anonymization techniques are not fool proof and can be broken with new information that an adversary may obtain or through weaknesses in the anonymization process. Federated learning techniques are costly to maintain and require consensus and centralization of models. This limits the ability to share data across organizations of diverse administrative and judiciary needs. One approach to address these challenges is the generation of synthetic data. Generating synthetic data provides a practical way to make data available while preserving privacy. Generative models produce data indistinguishable from real-world data while safeguarding privacy. Synthetic data offers cost-effective, scalable datasets that encourages data sharing. It reduces data privacy costs, fosters experimentation, enables collaboration, and expedites projects, seamlessly aligning with digital transformation goals. In this chapter, we review the work in this space and describe some of our own recent efforts.