Prediction of radionuclide diffusion enabled by missing data imputation and ensemble machine learning
摘要
Missing values in radionuclide diffusion datasets can undermine the predictive accuracy and robustness of the machine learning (ML) models. In this study, regression-based missing data imputation method using a light gradient boosting machine (LGBM) algorithm was employed to impute more than 60% of the missing data, establishing a radionuclide diffusion dataset containing 16 input features and 813 instances. The effective diffusion coefficient (De) was predicted using ten ML models. The predictive accuracy of the ensemble meta-models, namely LGBM-extreme gradient boosting (XGB) and LGBM-categorical boosting (CatB), surpassed that of the other ML models, with R2 values of 0.94. The models were applied to predict the De values of EuEDTA− and HCrO4− in saturated compacted bentonites at compactions ranging from 1200 to 1800 kg/m3, which were measured using a through-diffusion method. The generalization ability of the LGBM-XGB model surpassed that of LGB-CatB in predicting the De of HCrO4−. Shapley additive explanations identified total porosity as the most significant influencing factor. Additionally, the partial dependence plot analysis technique yielded clearer results in the univariate correlation analysis. This study provides a regression imputation technique to refine radionuclide diffusion datasets, offering deeper insights into analyzing the diffusion mechanism of radionuclides and supporting the safety assessment of the geological disposal of high-level radioactive waste.