This study addresses an integrate Parallel Machine Scheduling and Capacitated Vehicle Routing problem. In this problem, jobs must be processed efficiently on parallel machines before being distributed to customers through a fleet of vehicles. The main objective of this research is to optimize the machine production scheduling and the vehicle route selection minimizing the job’s total weighted tardiness. Since the problem is NP-Hard, we introduce a novel approach using Hybrid Metaheuristics and Reinforcement Learning Algorithms, the PPO-VND, that combines strengths using a Variable Neighborhood Descent Algorithm (VND) and leverages using Proximal Policy Optimization (PPO). The PPO-VND method employs PPO to dynamically select the most effective local search strategy for each iteration within the VND algorithm, evaluated through computational experiments on several instances. The results demonstrate that the VND approach assisted by a Proximal Policy Optimization using a Reinforcement Learning Algorithm outperforms traditional methods such as the Random Variable Neighborhood Descent Algorithm (RVND) and the Mixed-Integer Linear Programming (MILP).

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Reinforcement Learning Method for Integrated Production Scheduling and Distribution

  • Matheus de Freitas Araujo,
  • Thiago Henrique Nogueira,
  • José Elias Claudio Arroyo,
  • Julio César Alves

摘要

This study addresses an integrate Parallel Machine Scheduling and Capacitated Vehicle Routing problem. In this problem, jobs must be processed efficiently on parallel machines before being distributed to customers through a fleet of vehicles. The main objective of this research is to optimize the machine production scheduling and the vehicle route selection minimizing the job’s total weighted tardiness. Since the problem is NP-Hard, we introduce a novel approach using Hybrid Metaheuristics and Reinforcement Learning Algorithms, the PPO-VND, that combines strengths using a Variable Neighborhood Descent Algorithm (VND) and leverages using Proximal Policy Optimization (PPO). The PPO-VND method employs PPO to dynamically select the most effective local search strategy for each iteration within the VND algorithm, evaluated through computational experiments on several instances. The results demonstrate that the VND approach assisted by a Proximal Policy Optimization using a Reinforcement Learning Algorithm outperforms traditional methods such as the Random Variable Neighborhood Descent Algorithm (RVND) and the Mixed-Integer Linear Programming (MILP).