Modeling driver scanpath is crucial for understanding where and how drivers allocate attention in the surrounding environment, offering valuable insights into the decision-making processes of drivers and ultimately contributing to the enhancement of self-driving safety. However, existing methods primarily predict scanpath based on saliency maps, overlooking the dynamic sequential nature of gaze behavior. This limitation hinders the uncovering of underlying reasoning strategies and decision-making processes. In this work, we make the first attempt to predict both spatial coordinates and temporal duration of driver scanpath oriented by various tasks. Our method utilizes an innovative inverse reinforcement learning (IRL) framework and incorporates a spatial-temporal generator trained adversarially with a discriminator. Particularly, arguing that the driver’s decision-making processes are based on various driving tasks, we explicitly integrate task textual features into the generator to predict driver temporal scanpath aligning with various driver behaviors. Moreover, we fuse visual features with task-oriented information to construct a comprehensive context. In extensive experiments, our proposed method achieves the best results on DADA-Diverse and exhibits high interpretability.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Task-Oriented Scanpath Prediction with Spatial-Temporal Information in Driving Scenarios

  • Zhixin Huang,
  • Yuchen Zhou,
  • Chao Gou

摘要

Modeling driver scanpath is crucial for understanding where and how drivers allocate attention in the surrounding environment, offering valuable insights into the decision-making processes of drivers and ultimately contributing to the enhancement of self-driving safety. However, existing methods primarily predict scanpath based on saliency maps, overlooking the dynamic sequential nature of gaze behavior. This limitation hinders the uncovering of underlying reasoning strategies and decision-making processes. In this work, we make the first attempt to predict both spatial coordinates and temporal duration of driver scanpath oriented by various tasks. Our method utilizes an innovative inverse reinforcement learning (IRL) framework and incorporates a spatial-temporal generator trained adversarially with a discriminator. Particularly, arguing that the driver’s decision-making processes are based on various driving tasks, we explicitly integrate task textual features into the generator to predict driver temporal scanpath aligning with various driver behaviors. Moreover, we fuse visual features with task-oriented information to construct a comprehensive context. In extensive experiments, our proposed method achieves the best results on DADA-Diverse and exhibits high interpretability.