Rethinking crash risk prediction: classification versus probabilistic perspectives
摘要
In road traffic safety research, crash data is a major source of empirical evidence, which often involves discrete outcomes like crash occurrence and injury severity. Classification and probabilistic prediction are two feasible frameworks for crash data analysis. This paper clarifies the conceptual distinction between them, emphasising the practical importance of probabilistic prediction. Rather than a discrete class label, probabilistic models produce the probability of crash occurrence or serious/fatal injuries, which could be interpreted as traffic safety risk with inherent uncertainty measurement. To establish a reliable analytical strategy for probabilistic crash risk prediction, a series of simulation experiments were conducted, and a case study of single-bicycle crash injury severity was carried out. In particular, the consequences of data balancing treatments in class imbalance settings were investigated with both simulated and real-world crash data. Our evidence consistently indicates that data balancing will distort the prediction of probabilities. In our case study of single-bicycle crash injury severity, predictive models trained on rebalanced data yielded heavily over-estimated risk of fatalities. Therefore, data balancing treatments are not recommended for developing crash risk prediction models. Practical suggestions are summarised in accordance with these results. First, treat crash risk prediction as a probabilistic problem rather than a pure classification one. Second, abandon data balancing in developing crash risk prediction models. Lastly, choose performance metrics for probabilistic prediction (i.e., proper scoring rules) instead of classification accuracy to evaluate and compare the predictive models.