<p>Type 2 diabetes (T2D) is a growing global health crisis, affecting over 537 million people as of 2021. Early prediction remains particularly challenging in low- and middle-income countries due to missing data, class imbalance, and population-specific risk factors. This study presents a four-stage predictive framework— Feature-Weighted Class-Adaptive Generative Imputation Network-Weighted Classifier Aggregation Ensemble (FW-CAGIN-WCAE)—designed to address these limitations. First, Zero-Threshold Feature Removal (ZTFR) is applied to eliminate low-quality variables. Second, missing values are imputed FW-CAGIN, a novel class-aware and feature-weighted GAN model that accounts for both class and feature importance. Third, a performance-weighted ensemble of 15 machine and deep learning algorithms is constructed. Finally, SHAP analysis is used to uncover population-specific risk indicators. The proposed method was evaluated on three benchmark datasets—PIDD, FHGDD, and BDD—and their combinations, using nested five-fold cross-validation. The model achieved a peak AUC of 0.936 ± 0.018 in PIDD-BDD combination and reduced the imputation mean absolute error (MAE) from 0.8028 to 0.0033. It also lowered AUC variability by 36.3% and improved the diagnostic odds ratio (DOR) to 68.4 ± 20.5. SHAP analysis identified as a key predictive feature across both Asian and European populations. These findings demonstrate that the proposed framework offers an accurate, interpretable, and population-sensitive solution for early T2D detection, especially in resource-limited healthcare settings.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An ethnic-sensitive hybrid framework for T2D prediction with explainable AI and weighted ensembles

  • Karlo Abnoosian,
  • Rahman Farnoosh,
  • Hamidreza Noushkaran,
  • Danial Javaheri,
  • Hasnat Azeez Shimal,
  • Abdullah Abdulamer Abdulkarem

摘要

Type 2 diabetes (T2D) is a growing global health crisis, affecting over 537 million people as of 2021. Early prediction remains particularly challenging in low- and middle-income countries due to missing data, class imbalance, and population-specific risk factors. This study presents a four-stage predictive framework— Feature-Weighted Class-Adaptive Generative Imputation Network-Weighted Classifier Aggregation Ensemble (FW-CAGIN-WCAE)—designed to address these limitations. First, Zero-Threshold Feature Removal (ZTFR) is applied to eliminate low-quality variables. Second, missing values are imputed FW-CAGIN, a novel class-aware and feature-weighted GAN model that accounts for both class and feature importance. Third, a performance-weighted ensemble of 15 machine and deep learning algorithms is constructed. Finally, SHAP analysis is used to uncover population-specific risk indicators. The proposed method was evaluated on three benchmark datasets—PIDD, FHGDD, and BDD—and their combinations, using nested five-fold cross-validation. The model achieved a peak AUC of 0.936 ± 0.018 in PIDD-BDD combination and reduced the imputation mean absolute error (MAE) from 0.8028 to 0.0033. It also lowered AUC variability by 36.3% and improved the diagnostic odds ratio (DOR) to 68.4 ± 20.5. SHAP analysis identified as a key predictive feature across both Asian and European populations. These findings demonstrate that the proposed framework offers an accurate, interpretable, and population-sensitive solution for early T2D detection, especially in resource-limited healthcare settings.