Enhancing Traffic Crash Analysis with Fine-Tuned LLMs: A Social Media-Based Approach
摘要
In today’s big data era, social media platforms have emerged as vital sensors for monitoring real-time traffic incidents. Unlike traditional approaches that typically focus on multi-class classification to categorize data, this paper presents an innovative framework utilizing Large Language Models (LLMs) in a multitask learning approach, enabling the simultaneous extraction of diverse and detailed RTC-related information from the extensive and often noisy data found on Twitter. Initially, the GPT-3.5 model is utilized to extract six key features from reported Road Traffic Crashes (RTCs) related tweets. These features are then used to fine-tune a GPT-2 multitask classification model. This method enhances the ability to glean nuanced insights related to traffic incidents. Our advanced multitask framework significantly improves the detection and contextual understanding of RTCs. The fine-tuned GPT-2 model, trained on all tasks simultaneously, consistently outperformed other baseline models, including Logistic Regression, XGBoost, and AdaBoost, which were individually trained on each classification task. This demonstrates the model’s superior ability to classify and extract detailed information efficiently, thereby enhancing real-time monitoring and response to RTCs. The results underscore the potential of LLMs in revolutionizing traffic incident detection and analysis by leveraging the rich, real-time data available on social media.