Advancements and challenges in privacy-preserving split learning: experimental findings and future directions
摘要
Machine Learning (ML) has been extensively applied with remarkable success in various fields. However, training ML models using personal and/or private data has inadvertently revealed sensitive information and caused serious privacy risks. Consequently, Distributed Collaborative Machine Learning (DCML) has emerged as a promising approach to address such risks. Typically, DCML enables multiple entities to collaborate for training ML models with no disclosure of sensitive information. One of the most recent DCML techniques, namely Split Learning (SL), partitions the neural networks into a client and a server sub-network. Specifically, SL enables different entities to collaboratively train ML models and transmit the extracted features instead of sensitive raw data. Although SL prevents sending raw data to the server, the learned features may still reveal sensitive users’ information. This limitation has promoted Privacy-Preserving Split Learning (PPSL) which focuses on reinforcing data privacy within SL frameworks. This article represents a comprehensive survey on PPSL studies. It depicts relevant works published between 2018 and 2024. Moreover, the main objective of this research consists in identifying the key contributions of the state-of-the-art PPSL methods, in addition to investigating and comparing their performances. Accordingly, this survey supports the scientific community’s effort to bridge the research gaps and address challenges relevant to PPSL in resource-constrained environments. Furthermore, the findings of this survey highlight the importance of developing adaptive privacy protection techniques to brace future advancements in the field and strike a balance between data privacy and utility in the context of privacy-preserving split learning.