Fusion Framework for Accident Anticipation and Incident Detection in Dashcam Videos
摘要
This study introduces the Spatial–Temporal Attention (STA) network for early accident prediction from dashcam recordings. The primary aim is to leverage advancements in artificial intelligence and sensor technologies to enhance road safety. The research question focuses on whether the STA network can effectively identify crucial temporal segments and spatial regions in videos to predict accidents ahead of time. The STA network integrates a Temporal Attention (TA) component to discern significant temporal segments and a Spatial Attention (SA) component to concentrate on relevant spatial areas within video frames. Additionally, an attention module and Gated Recurrent Network (GRN) are employed to estimate the likelihood of future accidents. The network's performance is evaluated on three benchmark datasets, comparing it with other state-of-the-art methods in terms of average precision. A technique for fusing estimation scores from complementary models is also proposed and tested for improving prediction accuracy. The findings indicate that the STA network outperforms existing methods in early accident prediction, as demonstrated by higher average precision scores across the benchmark datasets. The incorporation of both temporal and spatial attention mechanisms allows the network to effectively identify discriminative segments and regions within video sequences, leading to more accurate predictions of potential accidents. Furthermore, the fusion technique enhances prediction performance, highlighting the efficacy of combining multiple models for improved accuracy. By leveraging spatial and temporal attention mechanisms, along with GRN and fusion techniques, the network demonstrates superior performance compared to existing methods. This advancement holds significant potential for enhancing road safety through proactive accident prevention measures.