There has been increasing interest in investigating the interplay of causality and RL in the recent years. The synergy of the two becomes understandable by treating the task of learning the action values Q(s, a) in RL as estimating the long-term counter-factual effects of applying the different actions to a given current state, to which knowing the corresponding causal structure in the environment is highly helpful.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Related Methods

  • Zhiwei (Tony) Qin,
  • Xiaocheng Tang,
  • Qingyang Li,
  • Hongtu Zhu,
  • Jieping Ye

摘要

There has been increasing interest in investigating the interplay of causality and RL in the recent years. The synergy of the two becomes understandable by treating the task of learning the action values Q(s, a) in RL as estimating the long-term counter-factual effects of applying the different actions to a given current state, to which knowing the corresponding causal structure in the environment is highly helpful.