Computer Vision and Deep Learning-Based Framework for Pedestrian Behavior Prediction Using Spatio-Temporal Features
摘要
This paper proposes a novel computer vision and deep learning based framework for understanding pedestrian behavior, focusing on road-crossing scenarios. Our approach utilizes input from dash-cam to analyze and interpret pedestrian actions, which is crucial for ensuring safety, especially for vulnerable road users. Estimating the road-crossing intentions of pedestrians is vital for enabling Advanced Driver Assistance Systems (ADAS) to make informed decisions and prevent potential collisions. Our framework detects and tracks pedestrians to extract spatio-temporal features and further, detects the pose to understand the changes in pose. Our proposed framework uses these features across the frames to predict whether a pedestrian will cross the road or not in advance before the actual action takes place. This information assists drivers to take informed decisions to avoid possible collisions. We have utilized the JAAD Dataset to perform the experiments and also collected some data from unstructured traffic environments to address the challenges in these scenarios. Experimental results demonstrates the robustness of our proposed framework in predicting pedestrian crossing behavior, thereby contributing to enhanced road safety.