Exploring Heterogeneity in Federated Learning
摘要
Federated Learning (FL) has emerged as a promising solution to address privacy concerns and data availability challenges in traditional machine learning (ML) by enabling collaborative model training across decentralized devices. Unlike conventional ML approaches that require centralized data aggregation, FL allows devices to train models locally and share only aggregated updates, preserving user privacy and reducing data transmission costs. However, FL faces significant challenges due to heterogeneity in various dimensions, including environmental factors, statistical data distributions, and model capabilities. This survey provides a comprehensive analysis of these challenges and recent advancements aimed at overcoming them. We explore environmental heterogeneity, which encompasses disparities in computational resources and network conditions, and discuss adaptive techniques to ensure fair participation and efficient model aggregation. Statistical heterogeneity, arising from non-independent and identically distributed (non-IID) data distributions across clients, is addressed through personalized and cluster-based approaches that enhance algorithm convergence and global model performance. In addition, we examine model heterogeneity, which requires resource-aware strategies to balance efficiency and robustness. Furthermore, the survey delves into the emerging paradigm of fully decentralized FL, leveraging technologies such as blockchain and topology optimization to enhance system scalability and resilience. Our findings highlight the importance of tackling heterogeneity to enable FL’s full potential in real-world applications. Future research should focus on scalable and privacy-preserving solutions to promote inclusive and efficient FL ecosystems in diverse environments.