<p>Vehicular edge computing enables vehicles to offload compute-intensive tasks to nearby edge servers integrated with roadside units to enhance performance and conserve battery power. In practical environments, vehicles often experience intrinsic task dependencies, stringent deadlines, and speeds and locations within highly dynamic networks, which is still a significant challenge in vehicular communications. Numerical optimization, game theory, and heuristic schemes struggle to meet these diverse and dynamic requirements in optimizing the computation offloading decisions. To this end, we propose a Proximal Policy Optimization (PPO)-based algorithm that optimally manages offloading decisions by employing a policy gradient approach with surrogate clipping to ensure stable and reliable updates. This is crucial in dynamic vehicular networks, where policy updates can significantly impact system performance. Furthermore, we use Generalized Advantage Estimation to further enhance the stability and efficiency by accurately estimating advantages over multiple steps. Our proposed PPO algorithm effectively balances exploration and exploitation and yields optimal offloading decisions while minimizing delays and dropped task ratio. Extensive experiments validate our approach’s efficacy, demonstrating substantial improvements in task completion rates and minimizing delays while meeting the requirements of intrinsic task dependencies and stringent deadlines in dynamic vehicular setups. For example, results show that the proposed method surpasses DQN by <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11227_2025_7009_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="51" /> </InlineMediaObject> <EquationSource Format="TEX">\(13.85\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>13.85</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation>, DDQN by <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11227_2025_7009_Article_IEq2.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="51" /> </InlineMediaObject> <EquationSource Format="TEX">\(11.24\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>11.24</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation>, DRQN by <InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11227_2025_7009_Article_IEq3.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="51" /> </InlineMediaObject> <EquationSource Format="TEX">\(19.48\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>19.48</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation>, and DA-TODDPG by <InlineEquation ID="IEq4"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11227_2025_7009_Article_IEq4.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="43" /> </InlineMediaObject> <EquationSource Format="TEX">\(8.81\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>8.81</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation> in terms of total delay, as well as achieving improvements of approximately <InlineEquation ID="IEq5"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11227_2025_7009_Article_IEq5.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="31" /> </InlineMediaObject> <EquationSource Format="TEX">\(30\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>30</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation> over DQN, <InlineEquation ID="IEq6"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11227_2025_7009_Article_IEq6.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="51" /> </InlineMediaObject> <EquationSource Format="TEX">\(17.65\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>17.65</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation> over DDQN, <InlineEquation ID="IEq7"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11227_2025_7009_Article_IEq7.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="43" /> </InlineMediaObject> <EquationSource Format="TEX">\(9.68\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>9.68</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation> over DA-TODDPG, and <InlineEquation ID="IEq8"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11227_2025_7009_Article_IEq8.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="31" /> </InlineMediaObject> <EquationSource Format="TEX">\(44\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>44</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation> over DRQN in terms of dropped task ratios at an RSU capability of <InlineEquation ID="IEq9"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11227_2025_7009_Article_IEq9.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="60" /> </InlineMediaObject> <EquationSource Format="TEX">\(140 \, \text {GHz}\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>140</mn> <mspace width="0.166667em" /> <mtext>GHz</mtext> </mrow> </math></EquationSource> </InlineEquation>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Computation offloading in vehicular communications using PPO-based deep reinforcement learning

  • Ehzaz Mustafa,
  • Junaid Shuja,
  • Faisal Rehman,
  • Abdallah Namoun,
  • Muhammad Bilal,
  • Adeel Iqbal

摘要

Vehicular edge computing enables vehicles to offload compute-intensive tasks to nearby edge servers integrated with roadside units to enhance performance and conserve battery power. In practical environments, vehicles often experience intrinsic task dependencies, stringent deadlines, and speeds and locations within highly dynamic networks, which is still a significant challenge in vehicular communications. Numerical optimization, game theory, and heuristic schemes struggle to meet these diverse and dynamic requirements in optimizing the computation offloading decisions. To this end, we propose a Proximal Policy Optimization (PPO)-based algorithm that optimally manages offloading decisions by employing a policy gradient approach with surrogate clipping to ensure stable and reliable updates. This is crucial in dynamic vehicular networks, where policy updates can significantly impact system performance. Furthermore, we use Generalized Advantage Estimation to further enhance the stability and efficiency by accurately estimating advantages over multiple steps. Our proposed PPO algorithm effectively balances exploration and exploitation and yields optimal offloading decisions while minimizing delays and dropped task ratio. Extensive experiments validate our approach’s efficacy, demonstrating substantial improvements in task completion rates and minimizing delays while meeting the requirements of intrinsic task dependencies and stringent deadlines in dynamic vehicular setups. For example, results show that the proposed method surpasses DQN by \(13.85\%\) 13.85 % , DDQN by \(11.24\%\) 11.24 % , DRQN by \(19.48\%\) 19.48 % , and DA-TODDPG by \(8.81\%\) 8.81 % in terms of total delay, as well as achieving improvements of approximately \(30\%\) 30 % over DQN, \(17.65\%\) 17.65 % over DDQN, \(9.68\%\) 9.68 % over DA-TODDPG, and \(44\%\) 44 % over DRQN in terms of dropped task ratios at an RSU capability of \(140 \, \text {GHz}\) 140 GHz .