This study deals with transfer reinforcement learning, i.e., reinforcement learning using knowledge acquired in one learning process (source task) in other learning processes (target tasks). The knowledge improves performance in the early stages of learning, but it may sometimes be detrimental, which is called negative transfer. To avoid the negative transfer, we propose a method in which the agent itself measures an effect of the transferred knowledge on the current learning process, and if it is harmful, it discards the knowledge during learning. We conducted experiments in a multi-player game environment with independently learning agents, in which the agents first learned their policies in one map of the environment and after that they learned in other, randomly generated maps. The result shows that the proposal mitigated negative transfer more successfully than existing methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Discarding Erroneous Knowledge Online in Transfer Reinforcement Learning

  • Otoya Notsu,
  • Koichi Moriyama,
  • Kosuke Shima,
  • Tohgoroh Matsui,
  • Atsuko Mutoh,
  • Nobuhiro Inuzuka

摘要

This study deals with transfer reinforcement learning, i.e., reinforcement learning using knowledge acquired in one learning process (source task) in other learning processes (target tasks). The knowledge improves performance in the early stages of learning, but it may sometimes be detrimental, which is called negative transfer. To avoid the negative transfer, we propose a method in which the agent itself measures an effect of the transferred knowledge on the current learning process, and if it is harmful, it discards the knowledge during learning. We conducted experiments in a multi-player game environment with independently learning agents, in which the agents first learned their policies in one map of the environment and after that they learned in other, randomly generated maps. The result shows that the proposal mitigated negative transfer more successfully than existing methods.