Incentivizing inclusive contributions in model sharing markets
摘要
Data plays a crucial role in training contemporary AI models, but much of the available public data will be exhausted in a few years, directing the world’s attention toward the massive decentralized private data. However, the privacy-sensitive nature of raw data and lack of incentive mechanism prevent these valuable data from being fully exploited. Here we propose inclusive and incentivized personalized federated learning (iPFL), which incentivizes data holders with diverse purposes to collaboratively train personalized models without revealing raw data. iPFL constructs a model-sharing market by solving a graph-based training optimization and incorporates an incentive mechanism based on game theory principles. Theoretical analysis shows that iPFL adheres to two key incentive properties: individual rationality and Incentive compatibility. Empirical studies on eleven AI tasks (e.g., large language models’ instruction-following tasks) demonstrate that iPFL consistently achieves the highest economic utility, and better or comparable model performance compared to baseline methods.