<p>Efficient traffic signal control is critical for reducing urban congestion, yet traditional rule-based and machine learning approaches fail to adapt to dynamic conditions. Reinforcement Learning (RL) provides a data-driven alternative, and this study evaluates classical exploration strategies Epsilon-Greedy, Thompson Sampling, Upper Confidence Bound (UCB), and Softmax alongside hybrid variants including Softmax with Temperature Annealing (SMX-TA), Softmax with Entropy Regularization (SMX-ER), and their combinations with UCB and Thompson Sampling. Our framework is based on value-based RL (Double Deep Q-Networks), chosen for scalability and stability compared to on-policy methods such as PPO and A2C, and implemented with parallel simulation in SUMO and CityFlow across synthetic grids (<InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11227_2025_7892_Article_IEq1.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="39" /> </InlineMediaObject> <EquationSource Format="TEX">\(1\times 1\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>1</mn> <mo>×</mo> <mn>1</mn> </mrow> </math></EquationSource> </InlineEquation>, <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11227_2025_7892_Article_IEq2.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="39" /> </InlineMediaObject> <EquationSource Format="TEX">\(4\times 4\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>4</mn> <mo>×</mo> <mn>4</mn> </mrow> </math></EquationSource> </InlineEquation>, <InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11227_2025_7892_Article_IEq3.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="39" /> </InlineMediaObject> <EquationSource Format="TEX">\(6\times 6\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>6</mn> <mo>×</mo> <mn>6</mn> </mrow> </math></EquationSource> </InlineEquation>) and real-world datasets (Hyderabad, New York). Benchmarking against classical baselines (FixedTime, MaxPressure) and recent RL models (FRAP, CoLight, GCN) shows that the proposed SMX-ER+UCB consistently yields the lowest average travel times and robust adaptability across traffic conditions. An ablation study confirms the complementary benefits of entropy regularization and UCB, while integration into scalable controllers such as CoLight demonstrates network-level generalization. These results highlight hybrid exploration as an effective and practical strategy for real-time, city-wide traffic optimization. Such large-scale optimization requires high-performance computing (HPC) to support parallel simulations, GPU-accelerated training, and multi-agent coordination, underscoring the necessity of HPC frameworks for intelligent transportation systems.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Reinforcement learning for traffic signal control: advancing efficiency through hybrid exploration strategies

  • Saidulu Thadikamalla,
  • Piyush Joshi,
  • Deepak Gangadharan

摘要

Efficient traffic signal control is critical for reducing urban congestion, yet traditional rule-based and machine learning approaches fail to adapt to dynamic conditions. Reinforcement Learning (RL) provides a data-driven alternative, and this study evaluates classical exploration strategies Epsilon-Greedy, Thompson Sampling, Upper Confidence Bound (UCB), and Softmax alongside hybrid variants including Softmax with Temperature Annealing (SMX-TA), Softmax with Entropy Regularization (SMX-ER), and their combinations with UCB and Thompson Sampling. Our framework is based on value-based RL (Double Deep Q-Networks), chosen for scalability and stability compared to on-policy methods such as PPO and A2C, and implemented with parallel simulation in SUMO and CityFlow across synthetic grids ( \(1\times 1\) 1 × 1 , \(4\times 4\) 4 × 4 , \(6\times 6\) 6 × 6 ) and real-world datasets (Hyderabad, New York). Benchmarking against classical baselines (FixedTime, MaxPressure) and recent RL models (FRAP, CoLight, GCN) shows that the proposed SMX-ER+UCB consistently yields the lowest average travel times and robust adaptability across traffic conditions. An ablation study confirms the complementary benefits of entropy regularization and UCB, while integration into scalable controllers such as CoLight demonstrates network-level generalization. These results highlight hybrid exploration as an effective and practical strategy for real-time, city-wide traffic optimization. Such large-scale optimization requires high-performance computing (HPC) to support parallel simulations, GPU-accelerated training, and multi-agent coordination, underscoring the necessity of HPC frameworks for intelligent transportation systems.