To effectively simulate crowd behavior, understanding the decision-making processes of pedestrians is paramount. This paper proposes a novel method for deducing pedestrians’ decision-making by conceptualizing the sum of their reward function along the trajectory as a utility function. While inverse reinforcement learning has been successfully used to retrieve the reward function of pedestrians, the outcome of advanced training algorithms is in the format of neural networks. Due to the black-box nature of neural networks, model interpretability methods are utilized to extract attributions of each input feature. This paper introduces a coupled method of inverse reinforcement learning and model interpretability to infer pedestrians’ decision-making on urban sidewalks based on the trajectory data collected in previous experiments. Furthermore, a preliminary test in a classical reinforcement learning environment cart pole is included to demonstrate the viability of the proposed method.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Inferring Pedestrian Decision-Making Through Inverse Reinforcement Learning

  • Xiangmin Yang,
  • Liu Yang,
  • Arnab Majumdar,
  • Washington Ochieng

摘要

To effectively simulate crowd behavior, understanding the decision-making processes of pedestrians is paramount. This paper proposes a novel method for deducing pedestrians’ decision-making by conceptualizing the sum of their reward function along the trajectory as a utility function. While inverse reinforcement learning has been successfully used to retrieve the reward function of pedestrians, the outcome of advanced training algorithms is in the format of neural networks. Due to the black-box nature of neural networks, model interpretability methods are utilized to extract attributions of each input feature. This paper introduces a coupled method of inverse reinforcement learning and model interpretability to infer pedestrians’ decision-making on urban sidewalks based on the trajectory data collected in previous experiments. Furthermore, a preliminary test in a classical reinforcement learning environment cart pole is included to demonstrate the viability of the proposed method.