Reinforcement learning for traffic signal control: advancing efficiency through hybrid exploration strategies
摘要
Efficient traffic signal control is critical for reducing urban congestion, yet traditional rule-based and machine learning approaches fail to adapt to dynamic conditions. Reinforcement Learning (RL) provides a data-driven alternative, and this study evaluates classical exploration strategies Epsilon-Greedy, Thompson Sampling, Upper Confidence Bound (UCB), and Softmax alongside hybrid variants including Softmax with Temperature Annealing (SMX-TA), Softmax with Entropy Regularization (SMX-ER), and their combinations with UCB and Thompson Sampling. Our framework is based on value-based RL (Double Deep Q-Networks), chosen for scalability and stability compared to on-policy methods such as PPO and A2C, and implemented with parallel simulation in SUMO and CityFlow across synthetic grids (