Optimizing LoRaWAN performance through Reinforcement Q-Convolutional Deterministic Policy Gradient: a comprehensive approach to efficient resource allocation
摘要
Low-power wide-area networks (LPWANs) is recognized as a leading technology for low-cost and energy-efficient communication in dense networks. However, the LoRa nodes face challenges in terms of data collection and energy consumption. To address these issues, we propose a Reinforcement Q-Convolutional Deterministic Policy Gradient (RQCDPG) method to optimize the physical layer (PHY) transmission parameters in LoRa networks. This advanced Reinforcement Learning (RL) technique combines the strengths of Q-learning and deterministic policy gradients, enhanced with convolutional layers to learn spatial and topological patterns from state features such as channel gain, task queue length, and residual capacity. The critic network estimates Q-values based on a reward function defined by energy consumption and transmission delay, while the actor network determines optimal continuous valued actions for PHY parameter selection. Unlike traditional methods, the proposed model operates only at the coordinator node, eliminating computational overhead at low-power end devices. Simulation results demonstrate that RQCDPG significantly improves Packet Delivery Ratio (PDR), energy efficiency, and latency compared to existing RL-based methods such as DDPG and ADR. This makes RQCDPG a scalable and practical solution for real-time and resource-constrained LoRaWAN scenarios.