Hierarchical Mean-Field Theory-based Off-Policy GRPO for Federated Edge Learning in Resource-Constrained Edge Computing
摘要
With the popularization of Internet of Things devices, the volume of data generated at the network edge has grown explosively. As a distributed machine learning solution, Federated Edge Learning (FEL) in mobile edge computing (MEC) provides an effective way to protect privacy by training locally and sharing only model parameters rather than the original data. However, in practical applications, FEL faces two core challenges: First, the data heterogeneity of each edge device can lead to deviations in the global model; The second is how to design an effective incentive mechanism to encourage nodes with limited resources to continuously participate in training. Our aims to simultaneously address these two major challenges by optimizing the local training and global aggregation processes to comprehensively enhance the efficiency and performance of FEL. For this reason, we propose a hybrid model. Firstly, the Stackelberg Stackelberg game model is adopted to describe the relationship between aggregators and edge devices. Meanwhile, the existence of Nash equilibrium is theoretically proved to ensure the stability of the model. Secondly, we propose a novel algorithm named Group Relative Policy Optimization Based on Hierarchical Mean-Field Theory (OGRPO-HMF), which can jointly optimize the local training of nodes and the global model aggregation of servers. We validate the effectiveness and generality of our approach through extensive experimentation on various FEL tasks, showcasing significant performance gains. Extensive experiments on benchmark FEL datasets demonstrate the superior performance of our proposed algorithm, improving the global test accuracy by up to