Optimizing Landmark Graphs in DHRL: A Dual Approach of Attention and Weighted Sampling
摘要
Graph-structured Deep Hierarchical Reinforcement Learning frameworks have made notable progress by modeling states, actions, and subtasks as graph nodes and edges, enabling rich semantic reasoning and capturing task dependencies. However, existing methods often suffer from inefficient experience replay and inaccurate subgoal relabeling, especially in sparse-reward and high-dimensional environments. This paper proposes a hybrid prioritized sampling mechanism based on state-goal distance and a temperature-regulated weighting strategy, along with an attention-driven subgoal refinement module integrated with a dual Q-network. Experiments across multiple benchmarks show that DAG-HRL achieves faster convergence and higher task success rates than prior methods.