Driver behavior significantly influences road safety, with human error contributing to the majority of traffic accidents globally. This paper presents an innovative AI-driven framework leveraging Visual Language Models (VLMs) to analyze driver behavior using ego vehicle video data from the Honda Research Institute Dataset, which includes naturalistic driving videos paired with Goal-Oriented Advice and Stimulus-Driven Advice. The framework utilizes the Gemini 2.0 Flash model in a zero-shot prompting setup, supported by YOLOv8 object detection model for traffic light recognition and fine-tuned MoViNet video classifier model for distinguishing between left turns, right turns, and straight movement, to generate scene descriptions and analyze driver behavior in complex traffic scenarios such as traffic lights, road signs, intersections, yielding, and turns. Results demonstrated a 77% accuracy rate in driver behavior analysis, while the reliability of zero-shot context learning remains inconsistent, highlighting areas for improvement. This study presents a preliminary analysis, which contributes to developing the full framework for safer roads by advancing automated driver assessment technologies. The framework can be incorporated into driver training programs, fleet management systems, or insurance risk assessment platforms.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Retrospective Evaluation of Driver Behavior Using Visual Language Models

  • Aya Hani,
  • Dina Omar,
  • Ziad Adel,
  • Abdallah A. Hassan,
  • Mohammed Elhenawy,
  • Huthaifa I. Ashqar,
  • Shadi Jaradat

摘要

Driver behavior significantly influences road safety, with human error contributing to the majority of traffic accidents globally. This paper presents an innovative AI-driven framework leveraging Visual Language Models (VLMs) to analyze driver behavior using ego vehicle video data from the Honda Research Institute Dataset, which includes naturalistic driving videos paired with Goal-Oriented Advice and Stimulus-Driven Advice. The framework utilizes the Gemini 2.0 Flash model in a zero-shot prompting setup, supported by YOLOv8 object detection model for traffic light recognition and fine-tuned MoViNet video classifier model for distinguishing between left turns, right turns, and straight movement, to generate scene descriptions and analyze driver behavior in complex traffic scenarios such as traffic lights, road signs, intersections, yielding, and turns. Results demonstrated a 77% accuracy rate in driver behavior analysis, while the reliability of zero-shot context learning remains inconsistent, highlighting areas for improvement. This study presents a preliminary analysis, which contributes to developing the full framework for safer roads by advancing automated driver assessment technologies. The framework can be incorporated into driver training programs, fleet management systems, or insurance risk assessment platforms.