Client heterogeneity and the Non-IID problem in federated learning systems have been major areas of focus for federated learning research. Convergence of model training is hindered by variations in device arithmetic settings and user data distributions. This paper presents a high-performance federated learning framework based on group normalization and local optimization gradient. The communication step is redesigned to upload the average value of the stochastic gradient instead of the final gradient at the end of the local optimization. In order to improve the model’s learning of the dataset representations and hasten convergence, a group normalization layer is simultaneously added to the local network and is not aggregated on the server side. By using the dynamic averaged gradient, FedGA successfully bridges the differences in the optimization process of different users and increases the model’s robustness. Convergence of the global model is also accelerated by the stronger local update. Our extensive testing on the image classification task shows that FedGA increases classification accuracy while preserving privacy; additionally, its convergence speed is nearly two times faster than the original FedAvg algorithm, and the accuracy is improved by approximately 7% on average when compared to FedAvg.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

FEDGA: FL with Dynamic Gradient and Group Normalization

  • Chuang Ma,
  • Xu Yang,
  • GuangXia Xu

摘要

Client heterogeneity and the Non-IID problem in federated learning systems have been major areas of focus for federated learning research. Convergence of model training is hindered by variations in device arithmetic settings and user data distributions. This paper presents a high-performance federated learning framework based on group normalization and local optimization gradient. The communication step is redesigned to upload the average value of the stochastic gradient instead of the final gradient at the end of the local optimization. In order to improve the model’s learning of the dataset representations and hasten convergence, a group normalization layer is simultaneously added to the local network and is not aggregated on the server side. By using the dynamic averaged gradient, FedGA successfully bridges the differences in the optimization process of different users and increases the model’s robustness. Convergence of the global model is also accelerated by the stronger local update. Our extensive testing on the image classification task shows that FedGA increases classification accuracy while preserving privacy; additionally, its convergence speed is nearly two times faster than the original FedAvg algorithm, and the accuracy is improved by approximately 7% on average when compared to FedAvg.