Cardiovascular disease remains a leading cause of mortality, emphasizing the need for accurate predictive models. The Cardiovascular Disease Dataset, consisting of 70,000 records, has been reported in the literature to present an underfitting problem, achieving only 73% accuracy. To address this, we enriched the dataset by calculating the Body Mass Index and generated a synthetic dataset of 2 million samples using a variational autoencoder generative adversarial network. Paired t-tests confirmed the synthetic data’s alignment with the original dataset’s statistical properties. A convolutional neural network was trained using 85% of the synthetic data for training and 15% for validation, reserving the original dataset for testing. The model achieved 92% accuracy on the original test data, with precision, recall, and F1 scores exceeding 90%. This study demonstrates that the model can effectively generalize from synthetic data to original data, addressing the underfitting issue and supporting the model’s robustness.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Improving Cardiovascular Risk Prediction Through Data Augmentation Using VAE-GAN and Deep Neural Networks

  • José L. López-Saynes,
  • Elías N. Escobar-Gómez,
  • Sergio F. Marroquín-Cano,
  • Eduardo Chandomi-Castellanos,
  • Sabino Velázquez-Trujillo,
  • Carlos A. Hernández-Gutiérrez,
  • Carlos V. de Coss-Pérez

摘要

Cardiovascular disease remains a leading cause of mortality, emphasizing the need for accurate predictive models. The Cardiovascular Disease Dataset, consisting of 70,000 records, has been reported in the literature to present an underfitting problem, achieving only 73% accuracy. To address this, we enriched the dataset by calculating the Body Mass Index and generated a synthetic dataset of 2 million samples using a variational autoencoder generative adversarial network. Paired t-tests confirmed the synthetic data’s alignment with the original dataset’s statistical properties. A convolutional neural network was trained using 85% of the synthetic data for training and 15% for validation, reserving the original dataset for testing. The model achieved 92% accuracy on the original test data, with precision, recall, and F1 scores exceeding 90%. This study demonstrates that the model can effectively generalize from synthetic data to original data, addressing the underfitting issue and supporting the model’s robustness.