In the era of data protection regulations like GDPR, safeguarding sensitive information has become paramount, prompting the exploration of synthetic data generation as a privacy-preserving alternative. Generative Adversarial Networks (GAN) and Variational Autoencoders (VAE), among other tools, have become popular for synthetic data generation. Despite their effectiveness, these models often carry the perception of being black boxes due to their complex learning mechanisms. Understanding the intricate behaviors of data within GAN or VAE poses a significant challenge, particularly with high-dimensional datasets. This is essential from privacy perspective as one can use synthetic data instead of original data and this can be considered as an alternative to anonymization. Our study aims to assess the distribution learning capabilities of synthetic data generators. Our methodology centers on artificially created datasets, such as swish roll and S-curve distributions, which offer easy visualization in \(\mathbb {R^n}\) space. Additionally, we evaluate point datasets containing discontinuous points to determine whether GAN and VAE comprehend the discontinuity behavior of datasets. By evaluating the data processed by GAN and VAE, we aim to reveal their learning capabilities and disentangle the complexities of synthetic data generation. Our research shifts the focus from real-world image datasets to artificially generated datasets, enabling exploration of commonly encountered distributions in low-dimensional spaces. Despite widespread recognition of GAN in image synthesis, achieving satisfactory results often requires employing numerous tricks due to training instability. We found that VAE exhibit a superior understanding of the underlying distribution of points in \(\mathbb {R^n}\) space compared to GAN. This inclination towards VAE arises from their more stable training process, inherent ability to capture latent structures within the data, and faster convergence compared to GAN.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Exploring Distribution Learning of Synthetic Data Generators for Manifolds

  • Sonakshi Garg,
  • Vicenç Torra

摘要

In the era of data protection regulations like GDPR, safeguarding sensitive information has become paramount, prompting the exploration of synthetic data generation as a privacy-preserving alternative. Generative Adversarial Networks (GAN) and Variational Autoencoders (VAE), among other tools, have become popular for synthetic data generation. Despite their effectiveness, these models often carry the perception of being black boxes due to their complex learning mechanisms. Understanding the intricate behaviors of data within GAN or VAE poses a significant challenge, particularly with high-dimensional datasets. This is essential from privacy perspective as one can use synthetic data instead of original data and this can be considered as an alternative to anonymization. Our study aims to assess the distribution learning capabilities of synthetic data generators. Our methodology centers on artificially created datasets, such as swish roll and S-curve distributions, which offer easy visualization in \(\mathbb {R^n}\) space. Additionally, we evaluate point datasets containing discontinuous points to determine whether GAN and VAE comprehend the discontinuity behavior of datasets. By evaluating the data processed by GAN and VAE, we aim to reveal their learning capabilities and disentangle the complexities of synthetic data generation. Our research shifts the focus from real-world image datasets to artificially generated datasets, enabling exploration of commonly encountered distributions in low-dimensional spaces. Despite widespread recognition of GAN in image synthesis, achieving satisfactory results often requires employing numerous tricks due to training instability. We found that VAE exhibit a superior understanding of the underlying distribution of points in \(\mathbb {R^n}\) space compared to GAN. This inclination towards VAE arises from their more stable training process, inherent ability to capture latent structures within the data, and faster convergence compared to GAN.