Research on Model-Free Reinforcement Learning Strategies for Airline Dynamic Pricing
摘要
In dynamic airline pricing, traditional operational optimization models, which rely on passenger demand forecasting, struggle to adapt to market fluctuations when demand is highly uncertain. To address this issue, this paper proposes a model for dynamic pricing based on model-free reinforcement learning (MFRL-DP), whose core lies in replacing demand forecasting with real-time market feedback, enabling continuous, self-adaptive learning of the strategy. Simulation experiments conducted using real airline data with limited attributes demonstrate that under this framework, commonly used methods such as DQN, A2C and PPO exhibit strong performance. Among them, PPO stands out for its strategy stability and overall flight revenue across different environments. This provides a novel, prediction-free, and highly adaptable approach to dynamic pricing.