Diffusion models for virtual populations and pharmacometric simulations
摘要
To evaluate diffusion model-based artificial intelligence approaches for generating virtual populations with physiological determinants of drug dosing (PDODD) and pharmacokinetic (PK) profiles. A denoising diffusion probabilistic model (DDPM) was applied to a 31-variable dataset of PDODD covariates (18 continuous, 13 binary) from the National Health and Nutrition Examination Survey and compared to a tabular variational autoencoder (TVAE). For nivolumab PK data (12,000 patients, 13 time points, 5 covariates), sequence-based diffusion model (SDM) and a time-aware diffusion model (TDM) with temporal self-attention were evaluated. The predictive performance of the TDM was evaluated by imputing masked time points. All models were trained and tested on 80%:20% partitions of the data using univariate, bivariate, and multivariate distributional similarity metrics. The diffusion model satisfactorily approximated the univariate distributions of continuous PDODD biomarkers (mean Kolmogorov-Smirnov D-statistic, KSD = 0.014), disease status frequencies (mean absolute error, MAE = 0.31%), and preserved bivariate correlations (MAE = 0.033). DDPM outperformed TVAE for categorical variables (0.31% vs. 1.07% MAE) and correlation (0.033 vs. 0.091 MAE). For nivolumab PK, SDM has KSD of 0.047 and a relative error of 1.36%. TDM accurately imputed missing PK timepoints (KSD = 0.014), reconstructing masked Day 1, Peak concentration (Cmax) Dose-9, and Terminal phase concentrations with MAE of 0.43%, 0.25%, and 0.76%, and correlations ≥ 0.999. Diffusion models demonstrated strong performance in generating cross-sectional PK covariate data and longitudinal PK profiles, capturing complex distributional and temporal dependencies. Diffusion-based approaches provide a flexible and robust framework for virtual simulations in pharmacometrics.