Same-day delivery problems are a class of stochastic decision making problems concerned with delivering orders placed dynamically by stochastic customers on the same day given a fleet of vehicles. We consider a variant where all orders have to be served with the objective to minimize a tardiness penalty function and where their spatiotemporal distribution is known. A well-known baseline approach to increase performance compared to myopic optimization is by sampling and optimizing scenarios in the short-horizon and deriving a consensus solution from the resulting plans. Its drawback is the computational effort required, which may not make it suitable for near real-time decision making. Extending recent methodology from the literature, we replace this online sampling by an offline training of a short-horizon value function using a neural network, which is then used in the online point-in-time optimization, combining current reward plus estimated future value of a solution candidate. In a first computational study on a single-vehicle instance class with unavoidable tardiness, we show that this leads to comparable performance as the sampling approach, while greatly reducing the online decision time.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Learning Value Functions for Same-Day Delivery Problems in the Tardiness Regime

  • Nikolaus Frohner,
  • Günther R. Raidl

摘要

Same-day delivery problems are a class of stochastic decision making problems concerned with delivering orders placed dynamically by stochastic customers on the same day given a fleet of vehicles. We consider a variant where all orders have to be served with the objective to minimize a tardiness penalty function and where their spatiotemporal distribution is known. A well-known baseline approach to increase performance compared to myopic optimization is by sampling and optimizing scenarios in the short-horizon and deriving a consensus solution from the resulting plans. Its drawback is the computational effort required, which may not make it suitable for near real-time decision making. Extending recent methodology from the literature, we replace this online sampling by an offline training of a short-horizon value function using a neural network, which is then used in the online point-in-time optimization, combining current reward plus estimated future value of a solution candidate. In a first computational study on a single-vehicle instance class with unavoidable tardiness, we show that this leads to comparable performance as the sampling approach, while greatly reducing the online decision time.