<p>In automated container terminals, improving the scheduling of automated stacking cranes (ASCs) is critical to optimizing operational efficiency. Since the twin ASCs cannot cross each other, a reasonable operation sequence to minimize interference and scheduling length (makespan) is needed. This problem is Nondeterministic Polynomial time (NP)-hard, and heuristic algorithms or optimization solvers cannot provide near-optimal solutions within a reasonable computational time. Therefore, a more efficient algorithm is necessary. In this paper, we construct a Markov Decision Process (MDP) model for the twin ASCs scheduling problem. To solve it, we propose a Proximal Policy Optimization algorithm with a Multi-head Self-attention mechanism (MHS-PPO), which is employed to capture complex task relationships under varying problem scales. Additionally, we propose a novel clock advancement method based on the future event list to power realistic task-execution simulations. Experimental results demonstrate that MHS-PPO outperforms traditional scheduling methods, revealing that the agent can effectively address operation sequencing decisions and crane interferences. The makespan obtained by MHS-PPO is approximately 10–15% better than the best-performing baseline method. In large-scale instances, it is still approximately 13.85% smaller than the Double Deep Q-network (DDQN) algorithm. The model exhibits excellent generalization capabilities in dynamic environments, effectively minimizing task waiting times. Moreover, it demonstrates robust adaptability across diverse task scales and distributions, showcasing its versatility and reliability in real-world applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A self-attention-based reinforcement learning approach for scheduling twin automated stacking cranes in container terminals

  • Liangcai Dong,
  • Yang Fan,
  • Zhennan Zhu,
  • Yuheng Liu

摘要

In automated container terminals, improving the scheduling of automated stacking cranes (ASCs) is critical to optimizing operational efficiency. Since the twin ASCs cannot cross each other, a reasonable operation sequence to minimize interference and scheduling length (makespan) is needed. This problem is Nondeterministic Polynomial time (NP)-hard, and heuristic algorithms or optimization solvers cannot provide near-optimal solutions within a reasonable computational time. Therefore, a more efficient algorithm is necessary. In this paper, we construct a Markov Decision Process (MDP) model for the twin ASCs scheduling problem. To solve it, we propose a Proximal Policy Optimization algorithm with a Multi-head Self-attention mechanism (MHS-PPO), which is employed to capture complex task relationships under varying problem scales. Additionally, we propose a novel clock advancement method based on the future event list to power realistic task-execution simulations. Experimental results demonstrate that MHS-PPO outperforms traditional scheduling methods, revealing that the agent can effectively address operation sequencing decisions and crane interferences. The makespan obtained by MHS-PPO is approximately 10–15% better than the best-performing baseline method. In large-scale instances, it is still approximately 13.85% smaller than the Double Deep Q-network (DDQN) algorithm. The model exhibits excellent generalization capabilities in dynamic environments, effectively minimizing task waiting times. Moreover, it demonstrates robust adaptability across diverse task scales and distributions, showcasing its versatility and reliability in real-world applications.