Traffic Accident Prediction Considering Imbalanced Data: An Interpretable Machine Learning Model
摘要
Traffic accident prediction is vital for enhancing road safety and mitigating accident risks. However, the inherent imbalance in accident data poses significant challenges to model performance, while existing models often neglect the temporal heterogeneity of holidays and peak hours and suffer from limited interpretability. This study presents a comprehensive traffic accident prediction framework to address these gaps. First, a novel data generation approach, incorporating gradient penalty terms into the Wasserstein generative adversarial network, is proposed to alleviate data imbalance and generate high-quality synthetic data. Second, time classification features, capturing the effects of time of day, holidays, and peak hours, are integrated into an enhanced extreme gradient boosting model for accident prediction. Lastly, the Shapley additive explanations method is applied to interpret the model results and uncover the nonlinear relationships between key factors and accident occurrence. Using accident and traffic data from the Xi’an ring expressway, the framework demonstrates superior performance over traditional methods in accuracy and reliability. Notably, average speed emerges as the most critical factor influencing accident risk, with a heightened risk observed when the mixing degree falls below 0.22 during peak hours or exceeds 0.7 universally. These findings provide actionable insights for proactive traffic safety management, enabling precise risk assessments and informed decision-making.