<p>High-performance computing (HPC) systems face increasing challenges in job scheduling due to the evolving complexity of computational tasks and the growing diversity and heterogeneity of resources. Job allocation is a critical aspect in contemporary HPC systems, due to compute nodes possessing an increased capacity in terms of physical resources and having the capability to execute multiple jobs simultaneously. However, job allocation is often overlooked in existing reinforcement learning (RL)-based schedulers that mainly focus on selecting suitable jobs from the job queue and leave allocation to overly simplistic policies, such as first-available allocation. The bin-packing nature at the node level of modern HPC necessitates more refined and intelligent allocation strategies. This paper introduces HeraSched, a novel hierarchical reinforcement learning (HRL)-based scheduler, adept at intelligent job selection without separate backfilling <i>and</i> heterogeneity-aware allocation, tailored for modern HPC environments. It efficiently manages diverse workloads across CPU and GPU cluster partitions. We evaluate HeraSched using real-world workloads, demonstrating significant improvements in reducing job waiting times and preventing job starvation compared to 27 scheduling combinations. In validation, the best maximum waiting time among compared methods is 78% higher than HeraSched’s result in overloaded CPU partitions. This performance demonstrates HeraSched’s ability to manage intensely stressed workloads and adapt to previously unseen, high-demand scenarios, thereby establishing a new standard in HPC job scheduling.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Optimizing HPC scheduling: a hierarchical reinforcement learning approach for intelligent job selection and allocation

  • Lingfei Wang,
  • Maria A. Rodriguez,
  • Nir Lipovetzky

摘要

High-performance computing (HPC) systems face increasing challenges in job scheduling due to the evolving complexity of computational tasks and the growing diversity and heterogeneity of resources. Job allocation is a critical aspect in contemporary HPC systems, due to compute nodes possessing an increased capacity in terms of physical resources and having the capability to execute multiple jobs simultaneously. However, job allocation is often overlooked in existing reinforcement learning (RL)-based schedulers that mainly focus on selecting suitable jobs from the job queue and leave allocation to overly simplistic policies, such as first-available allocation. The bin-packing nature at the node level of modern HPC necessitates more refined and intelligent allocation strategies. This paper introduces HeraSched, a novel hierarchical reinforcement learning (HRL)-based scheduler, adept at intelligent job selection without separate backfilling and heterogeneity-aware allocation, tailored for modern HPC environments. It efficiently manages diverse workloads across CPU and GPU cluster partitions. We evaluate HeraSched using real-world workloads, demonstrating significant improvements in reducing job waiting times and preventing job starvation compared to 27 scheduling combinations. In validation, the best maximum waiting time among compared methods is 78% higher than HeraSched’s result in overloaded CPU partitions. This performance demonstrates HeraSched’s ability to manage intensely stressed workloads and adapt to previously unseen, high-demand scenarios, thereby establishing a new standard in HPC job scheduling.