On the Low Rank Theory for the Data Shapley: Kernel Perspective
Abstract
Data Shapley offers a principled framework for data valuation, but remains costly to compute at scale. Although FreeShap avoids repeated model fine-tuning through kernel regression, evaluating many sampled coalitions remains computationally demanding. Low-rank kernel approximation offers a natural acceleration strategy, but its effect on individual Data Shapley values is poorly understood. We investigate how kernel approximation relates to prediction and Data Shapley. Our analysis shows that kernel approximation error controls errors in individual Data Shapley values, whereas predictive agreement alone does not guarantee Shapley agreement. Together with our experiments, these results indicate that eigendecomposition is more reliable than Nystr\"om approximation for Data Shapley. Guided by this analysis, we propose LRFShap, which accelerates coalition evaluation using shared low-dimensional kernel features. Across three downstream tasks, LRFShap achieves performance comparable to FreeShap at suitable ranks, with speedups of up to on data selection.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.