Related Methods
摘要
There has been increasing interest in investigating the interplay of causality and RL in the recent years. The synergy of the two becomes understandable by treating the task of learning the action values Q(s, a) in RL as estimating the long-term counter-factual effects of applying the different actions to a given current state, to which knowing the corresponding causal structure in the environment is highly helpful.