Harnessing machine learning algorithms for benchmarking deterioration prediction in civil infrastructure systems
摘要
The deterioration of civil infrastructure assets presents a serious global concern, affecting safety, functionality, and economic sustainability. Traditional statistical models, which often rely on linear assumptions and fixed deterioration rules, struggle to capture the complex patterns of asset degradation. This study conducts a comprehensive comparison of six machine learning algorithms i.e., multiple linear regression, decision tree regression, random forest regression, artificial neural network, extreme gradient boosting, and extra trees for predicting structural deterioration rates using real-world data from Indian bridges, roads, and pipelines. The dataset incorporates structural, environmental, operational, and maintenance-related variables. Models were rigorously trained using leakage-free cross-validation and evaluated using metrics such as coefficient of determination (R²), root mean squared error (RMSE), index of agreement (IOA), prediction interval (PI) coverage, and percentage of predictions within ± 20% of actual values (a20). Among all models, XGBoost demonstrated the highest predictive performance (R² = 0.87, RMSE = 0.93, PI coverage = 93.3%). Feature importance and interpretability were assessed using SHAP (SHapley Additive exPlanations), identifying age, chloride concentration, and traffic volume as the most influential predictors. The study provides a generalizable, interpretable, and uncertainty-aware framework for infrastructure asset management, offering practical guidance for data-driven maintenance planning and future extensions involving hybrid models and real-time sensor integration.