Collaborate Heterogeneous Resource Scheduling for AI Inference Based on Distributed Model Training in ECFN Enabled NG-RAN
摘要
As demand for artificial intelligence (AI) inference surges, traditional networks face significant challenges in computational and bandwidth resources. AI inference based on distributed model training can leverage parallel computing across multiple nodes to enhance training efficiency and model performance for subsequent real-time and accurate AI inference tasks. This paper mainly focuses on the collaborative scheduling issue of heterogeneous resources for AI inference services in Edge Computing First Network (ECFN) enabled NG-RAN. To address the issue, a mathematical model is proposed with the goal of minimizing transmission bandwidth usage rate, computing resource usage rate, backlog of AI inference subtasks, while maximizing the number of successfully sent AI inference subtasks. Moreover, we develop a heuristic algorithm called collaborate heterogeneous resource scheduling algorithm based on greedy strategy (CHRS-GS) in an efficient manner. Simulation results demonstrate that our proposed heuristic algorithm surpasses benchmark methods significantly achieving better performance in resource utilization, service success rates, and backlog minimization, reducing 21.4% joint optimization objectives, and approaching the optimal solution with an error of 3.1%.