<p>We address multi-period inventory decision-making using multisource multimodal data and propose a deep reinforcement learning (DRL) method—Word Embedding and Transformer-enhanced Twin Delayed Deep Deterministic Policy Gradient (WET-TD3). This method integrates multimodal environmental perception with policy optimization to produce end-to-end replenishment decisions for each period. First, we design multimodal feature-aware agent neural networks that incorporate word embeddings and Transformer modules to process structured demand-related features and unstructured customer reviews from multiple sources. This design constructs a state space responsive to dynamic markets. Second, we integrate into the TD3 algorithm a multimodal Actor-Critic architecture tailored for high-dimensional heterogeneous inputs. Additionally, we introduce delayed policy updates, experience replay, and exploration noise mechanisms to improve training stability. Experiments on real-world data show WET-TD3 outperforms benchmarks, reducing average cost by over 53.69%. It dynamically adjusts replenishment strategies based on the relative magnitudes of holding and underage costs, maintaining stable performance across cost structures. These results underscore the value of deeply integrating textual reviews and structured data, and demonstrate the DRL framework’s effectiveness for long-term optimization goals under demand uncertainty.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A multimodal deep reinforcement learning framework for multi-period inventory decision-making under demand uncertainty

  • Yu-Xin Tian,
  • Chuan Zhang

摘要

We address multi-period inventory decision-making using multisource multimodal data and propose a deep reinforcement learning (DRL) method—Word Embedding and Transformer-enhanced Twin Delayed Deep Deterministic Policy Gradient (WET-TD3). This method integrates multimodal environmental perception with policy optimization to produce end-to-end replenishment decisions for each period. First, we design multimodal feature-aware agent neural networks that incorporate word embeddings and Transformer modules to process structured demand-related features and unstructured customer reviews from multiple sources. This design constructs a state space responsive to dynamic markets. Second, we integrate into the TD3 algorithm a multimodal Actor-Critic architecture tailored for high-dimensional heterogeneous inputs. Additionally, we introduce delayed policy updates, experience replay, and exploration noise mechanisms to improve training stability. Experiments on real-world data show WET-TD3 outperforms benchmarks, reducing average cost by over 53.69%. It dynamically adjusts replenishment strategies based on the relative magnitudes of holding and underage costs, maintaining stable performance across cost structures. These results underscore the value of deeply integrating textual reviews and structured data, and demonstrate the DRL framework’s effectiveness for long-term optimization goals under demand uncertainty.