<p>The development of autonomous vehicles (AVs) increasingly relies on reinforcement learning (RL) within simulated environments. However, many existing frameworks report only end-of-training metrics, providing limited insight into how design choices influence learning outcomes. In this work, we introduce a unified simulation-based RL framework designed for both parking and driving tasks using the Unity engine. Our main contribution is a custom-built Adaptive Dual-Task Curriculum Scheduler (ADCS), which dynamically adjusts task difficulty to accelerate learning. We also develop a diagnostic logging pipeline that records per-step reward, loss, and entropy metrics, enabling fine-grained performance analysis. By integrating both PPO and DQN agents into this framework, we conduct systematic comparisons against multiple baselines and ablated variants. Our results demonstrate that ADCS, in combination with structured reward shaping, reduces training time by over 50% compared to flat curricula, while also improving training stability and policy generalization. Unlike prior work, our analysis explicitly quantifies the individual contributions of curriculum scheduling and reward engineering to learning speed and policy robustness. Finally, we propose a staged Sim-to-Real validation pathway—including domain-randomized simulation and closed-track testing—to address real-world transfer challenges. This work consolidates established RL techniques into a reproducible, statistically validated protocol that provides practical value for the safe and scalable training of autonomous vehicles.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comprehensive training and evaluation of autonomous vehicles using simulation-based reinforcement learning

  • Tamas Fulop,
  • Jozsef Katona

摘要

The development of autonomous vehicles (AVs) increasingly relies on reinforcement learning (RL) within simulated environments. However, many existing frameworks report only end-of-training metrics, providing limited insight into how design choices influence learning outcomes. In this work, we introduce a unified simulation-based RL framework designed for both parking and driving tasks using the Unity engine. Our main contribution is a custom-built Adaptive Dual-Task Curriculum Scheduler (ADCS), which dynamically adjusts task difficulty to accelerate learning. We also develop a diagnostic logging pipeline that records per-step reward, loss, and entropy metrics, enabling fine-grained performance analysis. By integrating both PPO and DQN agents into this framework, we conduct systematic comparisons against multiple baselines and ablated variants. Our results demonstrate that ADCS, in combination with structured reward shaping, reduces training time by over 50% compared to flat curricula, while also improving training stability and policy generalization. Unlike prior work, our analysis explicitly quantifies the individual contributions of curriculum scheduling and reward engineering to learning speed and policy robustness. Finally, we propose a staged Sim-to-Real validation pathway—including domain-randomized simulation and closed-track testing—to address real-world transfer challenges. This work consolidates established RL techniques into a reproducible, statistically validated protocol that provides practical value for the safe and scalable training of autonomous vehicles.