A reinforcement learning approach to dynamic portfolio optimization
摘要
This paper intends to bridge the gap between traditional and machine learning (ML) methods for dynamic portfolio optimization. We consider an investor who maximizes his utility from terminal wealth by dynamically allocating between risky and risk-free assets over time. Departing from the Deep Deterministic Policy Gradient (DDPG) algorithm by Lillicrap et al. (