<p>Urban rail transit train trajectory optimization is a constrained multi-objective sequential decision task plagued by sparse terminal rewards, coupled track gradient inertia, and multi-segment variable speed limits. Traditional Reinforcement Learning (RL) suffers from myopic policy optimization under sparse rewards, while vanilla Decision Transformer (DT) splits return-to-go (RTG), state and action into independent tokens, breaking the inherent Markov Decision Process (MDP) causal coupling and leading to severe attention imbalance. Meanwhile, existing sparse Transformers adopt single-granularity token-wise filtering which easily masks reward signals, and knowledge-enhanced Transformer schemes only inject static global physical bias without matching MDP temporal units, lacking explicit modeling of coupled track operation constraints. To address these gaps, this paper proposes a Domain-Knowledge integrated Block-Sparse Decision Transformer (DK-BSDT). Different from simple superposition of sparse attention and domain bias modules, our method delivers two fundamental architectural innovations: (1) A two-layer block-structured sparse self-attention mechanism reorganizes single-step RTG-state-action triplets into indivisible decision blocks, retaining intra-block MDP causality via full connection and designing multi-top-k weighted aggregation for inter-block multi-scale temporal feature screening to improve sparse reward credit assignment; (2) A dynamic domain knowledge bias matrix (DKBM) exclusively embedded into inter-block attention is constructed to synergistically encode long-term gradient inertial effects and hierarchical speed-limit constraints. Experimental results demonstrate that DK-BSDT achieves faster convergence, adapts to diverse typical line scenarios, and generates trajectories with lower energy consumption while meeting runtime deviation requirements. It exhibits stronger robustness in complex scenarios such as multi-speed limit transitions and steep gradients.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A decision transformer-based train trajectory optimization approach integrating domain knowledge and block-structured sparse self-attention

  • Ligang Cheng,
  • Jie Cao,
  • Ganxing Ouyang,
  • Shunpeng Hu

摘要

Urban rail transit train trajectory optimization is a constrained multi-objective sequential decision task plagued by sparse terminal rewards, coupled track gradient inertia, and multi-segment variable speed limits. Traditional Reinforcement Learning (RL) suffers from myopic policy optimization under sparse rewards, while vanilla Decision Transformer (DT) splits return-to-go (RTG), state and action into independent tokens, breaking the inherent Markov Decision Process (MDP) causal coupling and leading to severe attention imbalance. Meanwhile, existing sparse Transformers adopt single-granularity token-wise filtering which easily masks reward signals, and knowledge-enhanced Transformer schemes only inject static global physical bias without matching MDP temporal units, lacking explicit modeling of coupled track operation constraints. To address these gaps, this paper proposes a Domain-Knowledge integrated Block-Sparse Decision Transformer (DK-BSDT). Different from simple superposition of sparse attention and domain bias modules, our method delivers two fundamental architectural innovations: (1) A two-layer block-structured sparse self-attention mechanism reorganizes single-step RTG-state-action triplets into indivisible decision blocks, retaining intra-block MDP causality via full connection and designing multi-top-k weighted aggregation for inter-block multi-scale temporal feature screening to improve sparse reward credit assignment; (2) A dynamic domain knowledge bias matrix (DKBM) exclusively embedded into inter-block attention is constructed to synergistically encode long-term gradient inertial effects and hierarchical speed-limit constraints. Experimental results demonstrate that DK-BSDT achieves faster convergence, adapts to diverse typical line scenarios, and generates trajectories with lower energy consumption while meeting runtime deviation requirements. It exhibits stronger robustness in complex scenarios such as multi-speed limit transitions and steep gradients.