<p>This paper proposes a novel Dynamic Deep Reinforcement Learning (2DRL) framework that addresses the critical challenges of resource allocation and mode switching in Device-to-Device (D2D) networks operating under imperfect Channel State Information (CSI). Unlike traditional static schemes such as ARAMS, the proposed 2DRL dynamically learns optimal policies through real-time interaction with rapidly changing channel and interference conditions. Our design integrates advanced neural modules—Conv2D for spatial interference, LSTM for mobility awareness, Sinkhorn layers for resource allocation, and Gumbel-Softmax sampling for mode control—all constrained by practical 5 G NR requirements. Simulation results demonstrate that 2DRL achieves up to 22% higher throughput, maintains 85% performance retention under joint stressors, and reduces constraint violations by 6.8× compared to baselines. By significantly improving spectral efficiency, fairness, and energy usage, this work directly supports socially relevant goals such as reliable ultra-dense urban connectivity, smarter spectrum usage for next-generation IoT, and sustainable network operation for smart cities and 6 G ecosystems. The proposed framework lays the groundwork for future multi-agent reinforcement learning to further enhance scalable, low-latency, and resilient wireless infrastructures.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

2DRL: Cognitive D2D Control Under Imperfect CSI Via Adaptive Deep Reinforcement Learning

  • Panduranga Ravi Teja,
  • Krati Dubey,
  • Rishav Dubey

摘要

This paper proposes a novel Dynamic Deep Reinforcement Learning (2DRL) framework that addresses the critical challenges of resource allocation and mode switching in Device-to-Device (D2D) networks operating under imperfect Channel State Information (CSI). Unlike traditional static schemes such as ARAMS, the proposed 2DRL dynamically learns optimal policies through real-time interaction with rapidly changing channel and interference conditions. Our design integrates advanced neural modules—Conv2D for spatial interference, LSTM for mobility awareness, Sinkhorn layers for resource allocation, and Gumbel-Softmax sampling for mode control—all constrained by practical 5 G NR requirements. Simulation results demonstrate that 2DRL achieves up to 22% higher throughput, maintains 85% performance retention under joint stressors, and reduces constraint violations by 6.8× compared to baselines. By significantly improving spectral efficiency, fairness, and energy usage, this work directly supports socially relevant goals such as reliable ultra-dense urban connectivity, smarter spectrum usage for next-generation IoT, and sustainable network operation for smart cities and 6 G ecosystems. The proposed framework lays the groundwork for future multi-agent reinforcement learning to further enhance scalable, low-latency, and resilient wireless infrastructures.