<p>Synthetic data have been proposed to facilitate data sharing in privacy-sensitive contexts, including clinical trials. It remains unclear, however, how well original treatment effect estimates can be replicated in synthetic data analyses. Therefore, we synthesized and reanalyzed 128 treatment comparisons from 115 phase 3 randomized oncology trials using sixteen different generative models. Our findings demonstrate that careful methodological choices are essential for drawing valid statistical conclusions from synthetic data analyses. Naive analyses frequently yield falsely significant treatment effects, occurring in up to half of the trials created by deep generative models. Correcting standard errors to reflect the uncertainty inherent in synthetic data generation reduces these false positives, but primarily suffices for trials generated by parametric models. Although this correction entails some power loss, it can be mitigated by increasing the synthetic sample size. Thus, at present, large synthetic trials generated by parametric models and analyzed with corrected standard errors are more likely to preserve inferential utility. Advancing valid statistical inference from synthetic data created by deep generative models remains an important direction for future research.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Challenges of analyzing synthetic tabular data generated from 115 phase 3 oncology trials

  • Alexander Decruyenaere,
  • Christiaan Polet,
  • Johan Decruyenaere,
  • Heidelinde Dehaene,
  • Paloma Rabaey,
  • Thomas Demeester,
  • Stijn Vansteelandt,
  • Sylvie Rottey

摘要

Synthetic data have been proposed to facilitate data sharing in privacy-sensitive contexts, including clinical trials. It remains unclear, however, how well original treatment effect estimates can be replicated in synthetic data analyses. Therefore, we synthesized and reanalyzed 128 treatment comparisons from 115 phase 3 randomized oncology trials using sixteen different generative models. Our findings demonstrate that careful methodological choices are essential for drawing valid statistical conclusions from synthetic data analyses. Naive analyses frequently yield falsely significant treatment effects, occurring in up to half of the trials created by deep generative models. Correcting standard errors to reflect the uncertainty inherent in synthetic data generation reduces these false positives, but primarily suffices for trials generated by parametric models. Although this correction entails some power loss, it can be mitigated by increasing the synthetic sample size. Thus, at present, large synthetic trials generated by parametric models and analyzed with corrected standard errors are more likely to preserve inferential utility. Advancing valid statistical inference from synthetic data created by deep generative models remains an important direction for future research.