Secure Privacy-Preserving SMOTE for Vertical Federated Learning
摘要
In practical classification problems, the issue of sample imbalance is a pervasive challenge. The advent of federated learning, which involves the sharing of models among multiple participants without sharing data, has further complicated the handling of sample imbalance. This complexity is particularly pronounced in vertical federated learning, where a high degree of overlap in participant samples is required. The process of aligning samples while preserving privacy may result in a significant reduction in the available data samples, exacerbating the pre-existing imbalance issue. In the context of ensuring data privacy, we propose a secure privacy-preserving SMOTE (SP2-SMOTE) sampling method. It extends traditional SMOTE by allowing parties to independently generate synthetic samples without exposing the data, while effectively preventing unauthorized label inference through minority-class nearest neighbor interpolation. The evaluation of the imbalanced KEEL dataset, divided into two participants based on sample feature importance, demonstrates that SP2-SMOTE significantly improves the classification performance of vertical federated learning. These advances are validated by a series of metrics. This work offers a robust solution to the challenge of imbalanced data in vertical federated learning, rigorously preserving privacy for practical applications.