G-Value: Bridging the Gap Between LOO and Shapley Value Approaches for Data Valuation
摘要
As the amount of data used for machine learning training increases, estimating the contribution of individual data and selecting high-value data have become important issues. Many approaches have been proposed for data pricing and data selection, such as Leave-One-Out-based (LOO) score and a series of optimized approximation methods related to Shapley value. However, most methods assume that calculating marginal contributions and their expected values is effective, without analyzing the deeper theoretical differences between various methods. This paper, based on game interaction theory, attempts to design an optimized gradient-based contribution estimation method (referred to as G-Value) and conducts an initial exploration of its effectiveness from both theoretical and experimental perspectives, drawing conclusions that may be helpful for research in this field. In the experimental design, the G-Value algorithm is validated against a Monte Carlo approximation of the Shapley value. This research provides a unified framework to analyze the differences between various data analysis methods and offers guidance for future algorithm optimization.