<p>Distributed machine learning approaches are required when training data cannot be collected in a central location, due to storage, transmission or privacy/security constraints. An important task in any distributed machine learning context, and Federated Learning is no exception, is data value estimation or credit allocation, where the goal is to reward each participant proportionally to their contribution to the final performance of the machine learning model. However, all existing data value estimation techniques require that training be completed before the data values are obtained, and in this sense they can be considered as “a posteriori” approaches. Thus, all potential contributors must participate in the training process, regardless of the quality of their data or the final reward they can obtain. Here we present an “a priori” Shapley data value estimation technique in which, based on some statistical measures provided by the participants, the central counterpart or aggregator can obtain reasonably accurate data value estimates before actually starting the distributed learning process. To the best of our knowledge, this is the first “a priori” data value estimation approach proposed in the literature, and it can be used for the pre-selection of participants or to implement new pricing schemes. The introduced algorithms have been benchmarked using a variety of datasets and a logistic regression model, and we show that our “a priori” estimates are very accurate, compared to the centralized Shapley data values.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

“A priori” shapley data value estimation

  • Angel Navia-Vázquez,
  • Jesús Cid-Sueiro,
  • Manuel A. Vázquez

摘要

Distributed machine learning approaches are required when training data cannot be collected in a central location, due to storage, transmission or privacy/security constraints. An important task in any distributed machine learning context, and Federated Learning is no exception, is data value estimation or credit allocation, where the goal is to reward each participant proportionally to their contribution to the final performance of the machine learning model. However, all existing data value estimation techniques require that training be completed before the data values are obtained, and in this sense they can be considered as “a posteriori” approaches. Thus, all potential contributors must participate in the training process, regardless of the quality of their data or the final reward they can obtain. Here we present an “a priori” Shapley data value estimation technique in which, based on some statistical measures provided by the participants, the central counterpart or aggregator can obtain reasonably accurate data value estimates before actually starting the distributed learning process. To the best of our knowledge, this is the first “a priori” data value estimation approach proposed in the literature, and it can be used for the pre-selection of participants or to implement new pricing schemes. The introduced algorithms have been benchmarked using a variety of datasets and a logistic regression model, and we show that our “a priori” estimates are very accurate, compared to the centralized Shapley data values.